![]()
- GenBank is a major public nucleotide sequence database used by researchers around the world to share DNA sequence data with the scientific community. Researchers who generate new DNA sequences through methods such as Sanger sequencing, next-generation sequencing, genome sequencing, or targeted molecular studies may submit their sequence data to GenBank so that the information can become publicly available and identifiable through database accession numbers. Submitting sequences to GenBank is an important part of modern molecular biology and genomics because it supports data sharing, scientific transparency, reproducibility, and future research.
- The GenBank submission process involves more than simply uploading a DNA sequence. A successful submission generally requires researchers to prepare the nucleotide sequences, provide information about the biological source, describe relevant sequence features and annotations, check the data for errors, and submit the information through an appropriate NCBI submission system. The exact requirements depend on the type and complexity of the sequence data being submitted. Understanding the overall submission workflow before beginning can make the process considerably easier.
- Before submitting a sequence, researchers should first determine what type of sequence data they have generated. A single gene sequence, a collection of related DNA sequences, a complete genome, a transcriptome assembly, and Whole Genome Shotgun (WGS) data may follow different submission pathways or require different types of information. GenBank contains many categories of sequence data, so identifying the appropriate category is an important first step. The article Types of Sequence Data in GenBank provides a broader overview of these categories.
- The next step is to prepare the DNA sequence data. Sequence files should be checked carefully before submission because errors in nucleotide sequences can affect both the submission process and subsequent scientific analyses. Researchers should verify that the sequences contain valid nucleotide characters, that sequence orientations are appropriate, and that sequence lengths and identifiers are correct. When sequences originate from sequencing experiments, quality-control procedures should normally be completed before the data are prepared for submission.
- The biological source information associated with each sequence is also extremely important. GenBank records are not simply collections of nucleotide letters; they contain biological and contextual information that allows other researchers to understand what each sequence represents. Depending on the study, source information may include the organism, strain, isolate, tissue, specimen, developmental stage, geographic or environmental context, collection information, and other relevant characteristics. Accurate metadata increases the scientific usefulness of the submitted sequence.
- Researchers should also determine which biological features need to be represented in the sequence record. Depending on the sequence, these may include genes, coding sequences (CDS), messenger RNA, ribosomal RNA, transfer RNA, regulatory regions, or other relevant features. The locations of these features on the nucleotide sequence and their associated qualifiers may need to be provided as part of the submission. Understanding GenBank sequence annotation and the GenBank feature table is therefore useful when preparing more complex submissions.
- For coding sequences, additional information may be required to describe the coding region and its biological product. Researchers may need to identify the location of the CDS and provide appropriate information about the encoded product. For RNA sequences, genes and RNA features may need to be described according to the nature of the sequence. The level of annotation required depends on the sequence and the type of submission, so researchers should follow the current NCBI submission guidance applicable to their dataset.
- An important part of preparing a GenBank submission is ensuring that the submitted information accurately represents the underlying biological material. Sequence identifiers, organism names, source information, feature coordinates, and other metadata should be internally consistent. For example, the organism associated with a sequence should correspond to the biological sample from which the sequence was obtained. Incorrect metadata can make a sequence difficult to interpret and may require corrections after submission.
- Researchers should also consider whether the sequence is complete, partial, or otherwise limited in scope. A partial gene sequence should not be represented as a complete gene sequence, and an incomplete genomic sequence should not be described as a complete genome. Descriptions and annotations should accurately reflect what the experimental data support. Clear and accurate descriptions help downstream users select appropriate records for comparative analyses.
- NCBI provides several tools and submission pathways for depositing nucleotide sequence data. The appropriate tool depends on the type and scale of the submission. Smaller or relatively straightforward submissions may be prepared through web-based submission systems, whereas large or specialized datasets may require other approaches. Genome-scale and high-throughput projects can involve additional resources and relationships between sequence data, biological samples, projects, and genome assemblies.
- One commonly encountered submission resource is BankIt, a web-based system designed for submitting sequence data to GenBank. Other submission mechanisms are available for particular types of datasets and workflows. Researchers should select the submission method appropriate to their sequence type rather than assuming that every GenBank submission follows exactly the same procedure.
- During submission, researchers may be asked to provide information about the submitter and the sequence project. They may also need to provide publication information or indicate whether the sequence is associated with a manuscript. The information supplied during submission contributes to the metadata associated with the resulting GenBank record. Researchers should therefore ensure that names, affiliations, contact information, project descriptions, and publication details are accurate.
- Sequence annotation is another important component of submission. A researcher submitting a newly generated gene sequence may need to identify the gene and coding region, while a larger genomic submission may contain many genes and other sequence features. Annotation can range from relatively simple feature information to extensive genome-level annotation. The goal is to provide sufficient biological information for the sequence to be understood and used appropriately.
- Validation is an important stage before a submission is finalized. Submission systems can identify certain problems with sequence formatting, feature locations, qualifiers, taxonomy, or other required information. Researchers should carefully review validation messages and correct problems before completing the submission. A successful validation process does not necessarily guarantee that every biological interpretation is correct, so researchers should still manually review the sequence and metadata.
- Taxonomic information is particularly important for GenBank submissions. The organism associated with a sequence needs to be represented appropriately within the relevant NCBI taxonomy framework. Researchers working with unusual organisms, newly characterized organisms, environmental samples, or uncertain taxonomic identifications may need additional guidance. Taxonomic information should not be guessed simply to make a submission easier; it should reflect the available scientific evidence.
- For some projects, sequence submission is closely connected to other NCBI resources. Large datasets may involve BioProject records that describe the overall research project and BioSample records that describe biological samples. These resources provide additional context around sequence data and can help connect individual sequence records to broader studies. Understanding these relationships becomes increasingly important as the scale of a sequencing project increases.
- Genome projects require particular attention because a genome may involve many individual sequence records and may also be represented through an assembly resource. Whole Genome Shotgun submissions, for example, can involve large numbers of sequence records generated from fragmented genomic sequencing. Researchers should distinguish between submitting WGS sequence data and creating or submitting a genome assembly. These processes are related but are not necessarily identical.
- Transcriptome projects can involve another specialized category of sequence data. Transcriptome Shotgun Assembly (TSA) submissions represent assembled transcript sequences rather than raw sequencing reads. Researchers working with transcriptomic datasets should therefore understand the distinction between raw sequencing data, assembled transcripts, and the corresponding GenBank submission categories.
- After a submission has been reviewed and processed, GenBank assigns accession identifiers to accepted sequence records. An accession number provides a way to locate and reference the submitted sequence in the database. Researchers should retain these identifiers because they are important for publications, data management, collaboration, and future retrieval of the submitted records. When sequence versions are relevant, the accession.version identifier can provide additional precision.
- The accession number is not simply an administrative detail. Once a sequence becomes part of the public database, the identifier provides a persistent reference point for researchers who need to locate the record. A publication describing newly generated sequences can therefore cite the relevant accession numbers so that readers can access the underlying sequence data. This strengthens transparency and allows other researchers to use the data in subsequent analyses.
- Researchers should also understand that submitting a sequence to GenBank makes the data available as part of a public scientific database according to the applicable GenBank policies and release procedures. Before submission, investigators should make sure that they understand any institutional, ethical, contractual, or project-specific requirements governing the release of their sequence data. Particular care may be necessary when sequence information is associated with sensitive biological or human-related data.
- The timing of public release can also be relevant to researchers preparing manuscripts. Depending on the circumstances of a submission, researchers may need to coordinate sequence deposition with publication plans. Some journals require nucleotide sequence accession information as part of manuscript submission or publication. Researchers should therefore consider database submission early in the research and publication workflow rather than waiting until the final stages of manuscript preparation.
- A common mistake is submitting sequence data before adequately checking the underlying files and metadata. Other problems can include incorrect organism information, incomplete source metadata, incorrect feature coordinates, inappropriate feature qualifiers, inconsistent sequence identifiers, and inaccurate descriptions of partial or complete sequences. These issues can lead to delays, corrections, or records that are less useful to the scientific community.
- Another common problem is confusing sequence submission with sequence analysis. GenBank is primarily a public repository for sequence data and associated biological information. Researchers may perform extensive analyses before submission, including sequence quality assessment, assembly, annotation, alignment, phylogenetic analysis, or similarity searches. These analyses help determine what information should be submitted, but they are separate from the basic process of depositing the sequence data in GenBank.
- Good record keeping is essential throughout the submission process. Researchers should retain copies of the original sequence files, final submitted files, metadata, annotation information, validation messages, submission identifiers, correspondence, and accession numbers. Maintaining these records makes it easier to correct a submission if necessary and provides a clear data trail for future research.
- Once accession numbers have been assigned, researchers can use them to retrieve their records and verify that the publicly available information is correct. It is good practice to inspect the resulting GenBank record rather than assuming that the final database record exactly matches what was originally intended. Researchers can check the sequence, annotations, source information, references, and other record components after processing.
- The submission process is also closely related to the structure of a GenBank record. A submitted sequence ultimately becomes a structured database record containing information such as the definition, accession identifier, organism and source information, references, features, and nucleotide sequence. Understanding how a GenBank record is organized makes it easier to prepare appropriate information during submission and to recognize how the submitted data will appear to other researchers.
- For students and researchers submitting a small number of DNA sequences, the overall workflow can be summarized as follows: determine the appropriate sequence type, prepare and quality-check the nucleotide sequences, gather biological source information, prepare relevant annotations, select the appropriate NCBI submission system, enter or upload the required information, validate the submission, correct any problems, submit the data, and retain the resulting accession numbers. More complex projects may require additional project, sample, assembly, or annotation resources.
- For large-scale sequencing projects, submission planning should begin much earlier. Genome and transcriptome projects can generate enormous quantities of data, and organizing sequences, samples, metadata, annotations, and project information after the sequencing project is complete can be difficult. Establishing consistent identifiers and metadata structures from the beginning can make later GenBank submission substantially easier.
- GenBank submission also contributes to the broader international exchange of nucleotide sequence information. GenBank participates in the International Nucleotide Sequence Database Collaboration (INSDC) together with the European Nucleotide Archive (ENA) and the DNA Data Bank of Japan (DDBJ). Through this collaboration, nucleotide sequence information submitted to participating databases becomes part of an international system for sharing sequence data.
- Submitting DNA sequences to GenBank is therefore both a technical and scientific process. The technical side involves preparing sequence files, metadata, annotations, and submission information, while the scientific side involves ensuring that the submitted information accurately describes the biological material and experimental results. Careful preparation improves the quality of the resulting database record and makes the data more useful to researchers who may rely on it in future studies.
- Ultimately, a successful GenBank submission transforms newly generated DNA sequence data into a standardized, identifiable, and publicly accessible scientific resource. Whether the project involves a single gene, a collection of targeted sequences, a genome, or a large sequencing project, the same fundamental principles apply: prepare accurate sequences, provide appropriate biological context, validate the information carefully, follow the correct submission pathway, and preserve the resulting accession information. Detailed guides on the GenBank submission process, GenBank submission tools, and sequence preparation can then be used to explore each stage in greater depth.