![]()
- The GenBank submission process is the procedure researchers use to deposit newly generated nucleotide sequence data into GenBank, the public nucleotide sequence database maintained by the National Center for Biotechnology Information (NCBI). Submitting sequence data allows researchers to make their DNA and RNA sequences publicly available, obtain accession numbers, and provide other scientists with access to the underlying sequence information. For many research publications, sequence deposition is also an important part of scientific data sharing and reproducibility. NCBI currently provides a Submission Portal for GenBank sequence submissions, with workflows that vary according to the type of sequence being submitted.
- The first step in the GenBank submission process is determining whether GenBank is the appropriate destination for the data. Not every type of sequencing data should be submitted directly to GenBank. For example, raw unassembled next-generation sequencing reads are generally submitted to the Sequence Read Archive (SRA), while assembled transcriptome sequences may be submitted through the TSA workflow and complete or draft prokaryotic and eukaryotic genomes through the appropriate genome submission workflow. Choosing the correct NCBI resource before preparing the submission can prevent unnecessary reformatting and delays.
- For ordinary assembled nucleotide sequence submissions, researchers should first determine the biological nature of their sequences. These may include individual genes, coding regions, RNA sequences, organelle sequences, plasmids, viral sequences, environmental sequences, or other assembled nucleotide regions. The current Submission Portal GenBank provides organism- and sequence-specific workflows, including workflows for prokaryotes, eukaryotes, viruses, and synthetic constructs. Some sequence types receive automatic feature annotation, while others require the submitter to provide annotation.
- Before beginning a submission, researchers should prepare the information that will be required by the submission system. NCBI’s current checklist includes submitter contact information, author and publication information, sequencing technology, nucleotide sequences, biological source information, and feature annotation when required. Preparing these materials in advance makes the submission process considerably easier.
- The nucleotide sequences themselves are commonly prepared in FASTA format for submission through the current Submission Portal GenBank workflow. A FASTA file contains a definition line beginning with the greater-than symbol (>) followed by a sequence identifier and the nucleotide sequence. NCBI recommends keeping sequence identifiers short, simple, unique, and without spaces. For the current Submission Portal GenBank workflow, a single FASTA file can contain up to 3,000 sequences, although fewer sequences may be appropriate when the sequences are particularly long.
- Sequence quality should be checked before uploading the FASTA file. Researchers should verify that the sequences are correctly oriented, contain appropriate nucleotide characters, have the expected lengths, and represent the intended biological material. Sequencing errors, incorrect identifiers, duplicated sequences, accidental contamination, and improperly formatted files can cause problems during validation or downstream processing.
- Source information is another essential component of a GenBank submission. The source describes the biological or environmental material from which the sequence was obtained. Depending on the organism and sequence type, relevant information can include the organism name, strain, isolate, clone, specimen voucher, collection date, geographic location, host, isolation source, or other biological descriptors. NCBI checks organism names against its Taxonomy database, and additional information may be requested when an organism is not recognized.
- Researchers should pay particular attention to consistency between the sequence and its source metadata. The organism name, strain or isolate designation, collection information, and other descriptors should correspond to the actual biological sample. Incorrect or incomplete metadata can make an otherwise useful sequence difficult to interpret. For this reason, metadata preparation should be treated as an integral part of sequence submission rather than an administrative detail.
- The next major component is sequence annotation. GenBank features represent biological elements within nucleotide sequences, such as genes, coding regions, promoters, ribosomal RNAs, and other features. Depending on the submission type, features may be automatically annotated by NCBI or may need to be supplied by the researcher. For submissions requiring submitter-provided annotation, NCBI’s current Submission Portal supports several approaches, including entering information through the submission interface, uploading a five-column tab-delimited feature table, or providing a protein FASTA file for CDS annotation.
- Feature annotation should be reviewed carefully before submission. Feature coordinates must correspond to the actual nucleotide sequence, and qualifiers should accurately describe the biological feature. For coding sequences, researchers should check the coding region and its translation when appropriate. Incorrect feature coordinates or inconsistent annotations can generate validation errors or lead to inaccurate biological information in the resulting GenBank record.
- Taxonomy is another important part of the submission process. The scientific name associated with a sequence should reflect the organism from which the sequence was obtained and should be consistent with available taxonomic information. Researchers should not select an organism name merely because it appears similar to their sample. If the organism cannot be identified to the species level, an appropriate higher-level taxonomic designation may sometimes be used according to NCBI’s submission guidance.
- After the sequence, source information, and annotation have been prepared, the researcher can begin the submission through the appropriate NCBI submission system. As of 2026, the Submission Portal GenBank is the primary web-based route for most GenBank sequence submissions. NCBI has been transitioning functionality from BankIt to the Submission Portal, with BankIt being phased out during 2026. Researchers following older GenBank tutorials should therefore check the current NCBI submission guidance rather than relying on historical screenshots or instructions.
- The Submission Portal GenBank provides workflows tailored to different types of sequence data. Current workflows include prokaryote, eukaryote, virus, and synthetic construct submissions. This organism-specific structure is intended to guide submitters toward the appropriate metadata and annotation requirements. The portal also provides options for entering or uploading source information and reviewing the submission before processing.
- Once the appropriate workflow has been selected, the submitter provides the sequence files and associated information requested by the portal. The exact fields depend on the submission type. Researchers should complete each section carefully and avoid entering placeholder information that does not accurately describe the sequence or biological source.
- The submission system performs validation and processing after the data are submitted. NCBI describes a processing workflow that includes automated validation and manual review. A submission may enter an error state if information needs to be corrected, or a processing state while NCBI reviews the submission. After accession numbers have been assigned, the submission can proceed through additional processing before it becomes ready for public release.
- Validation errors should be treated as useful diagnostic information rather than simply as obstacles. They may indicate incorrect formatting, missing source information, invalid feature coordinates, taxonomy problems, inconsistent qualifiers, or other issues that need attention. Researchers should carefully read each error or warning, determine its cause, make the necessary correction, and review the submission again before finalizing it.
- One important distinction is the difference between errors and biological judgment. A submission system can identify many technical inconsistencies, but automated validation cannot replace scientific review. Researchers remain responsible for ensuring that the sequence, source metadata, annotation, organism identification, and other information accurately represent the underlying study.
- Once the submission has been processed and accession numbers have been assigned, researchers should record those accession numbers carefully. An accession number provides an identifier that allows the sequence record to be located and cited. GenBank notes that accession numbers can typically be assigned relatively quickly for standard submissions, although processing time can vary depending on submission type and complexity.
- Accession numbers are particularly important when a sequence is being submitted in connection with a manuscript. Many journals require nucleotide sequence data underlying a publication to be deposited in a public sequence database. GenBank participates in the International Nucleotide Sequence Database Collaboration with ENA and DDBJ, so researchers generally need to deposit their sequence data in only one of these collaborating repositories.
- Researchers should also understand that accession assignment and public release are related but distinct stages. GenBank can process sequence submissions and assign accession identifiers before the data are publicly released, allowing researchers to include accession information in manuscripts while coordinating data release with publication. The specific release arrangements depend on the submission and applicable policies.
- One common mistake is choosing the wrong submission pathway. A researcher may attempt to submit raw sequencing reads directly to GenBank when those data belong in SRA, or may attempt to submit a complete genome through a workflow intended for individual assembled sequences. NCBI’s Submission Portal is designed to direct users toward different resources depending on the data type, so researchers should determine whether their data are assembled sequences, raw reads, genome assemblies, transcriptome assemblies, or another category before starting.
- Another common mistake is confusing WGS data with a finished genome assembly. A WGS submission can contain multiple sequence pieces representing a genome that has not been completely resolved into single chromosome sequences. Genome submissions have additional requirements concerning sequence organization, chromosome or plasmid assignments, gaps, and related metadata. Researchers submitting genomes should therefore use the appropriate genome submission guidance rather than treating the submission as a collection of ordinary gene sequences.
- Submitting sequences that were not actually generated or determined by the submitting research group can also create problems. GenBank’s submission policies distinguish direct submissions from other types of sequence data, and standard submissions are intended for nucleotide sequences determined by or on behalf of the submitter. Researchers should therefore establish the provenance of the sequence data before submission.
- Another mistake is submitting consensus or artificially constructed sequence information without considering whether it meets GenBank’s submission criteria. Current GenBank guidance specifies categories of data that are not accepted as standard submissions, including certain noncontiguous sequences, primer sequences, protein sequences without an underlying nucleotide submission, mixed genomic/mRNA sequences, and sequences without an appropriate physical counterpart.
- Incomplete source metadata is another frequent problem. A sequence without sufficient information about the organism or biological source may be difficult for other researchers to interpret and may require additional information during submission processing. Researchers should collect metadata during the experimental stage rather than attempting to reconstruct it months or years later.
- Incorrect annotation is another important source of problems. A gene feature placed at the wrong coordinates, an incorrect coding sequence, or an inappropriate qualifier can affect how users interpret the sequence. Before submission, researchers should compare the annotation against the underlying sequence and confirm that the biological interpretation is supported by the available evidence.
- A further mistake is failing to distinguish partial sequences from complete sequences. If a sequence represents only part of a gene or genomic region, the record should accurately indicate its partial nature. Describing a partial sequence as complete can mislead researchers who later use the record for comparative analyses, primer design, phylogenetic studies, or genome annotation.
- Researchers should also avoid relying exclusively on old GenBank submission tutorials. The submission environment is evolving, and NCBI has introduced the Submission Portal GenBank as the newer unified workflow while transitioning away from BankIt. Older resources may still be useful for understanding concepts, but current NCBI instructions should be consulted for the actual submission interface and supported workflows.
- Another practical mistake is failing to keep copies of the submitted data. Researchers should retain the original sequence files, final FASTA files, metadata, feature tables, submission information, validation messages, and accession numbers. These records can be invaluable if corrections are required later or if collaborators need to reproduce the submission.
- After accession numbers have been assigned, researchers should inspect the resulting GenBank records when they become available. Checking the final record provides an opportunity to verify the sequence, organism information, annotations, references, and other metadata. If corrections are needed after accession assignment, researchers should use the appropriate GenBank update process rather than simply creating a duplicate new submission.
- Large sequencing projects require additional planning. Genome projects may involve BioProject and BioSample records, genome assemblies, WGS data, and extensive annotation. Transcriptome projects may involve SRA data together with TSA assemblies. Establishing project and sample identifiers early can make it easier to connect the resulting GenBank records to the broader study.
- Privacy and data-release considerations are also important. Researchers working with human-associated sequence data should understand the applicable GenBank policies and should not submit information that could reveal the identity of the source. NCBI’s current policy also states that, effective June 16, 2026, GenBank no longer accepts personal sequence data from private individuals, including certain personal human chromosome, genome, or mitochondrial DNA data.
- For a straightforward assembled DNA sequence, the overall workflow can be summarized as follows: determine that GenBank is the appropriate repository, prepare and quality-check the sequence, create the required FASTA file, collect accurate source metadata, prepare feature annotations when necessary, select the appropriate Submission Portal GenBank workflow, upload or enter the required information, review and validate the submission, correct any errors, submit the data for processing, obtain the accession number, and verify the final GenBank record.
- The exact workflow becomes more specialized as the dataset becomes more complex. A single gene sequence may require relatively limited information, whereas a genome, viral dataset, environmental project, or large batch of sequences can involve additional metadata, annotation, project identifiers, and submission requirements. Researchers should therefore avoid assuming that a submission procedure designed for one sequence type will apply unchanged to another.
- The most important principle is to treat GenBank submission as part of the scientific data-management workflow rather than as an administrative task performed immediately before publication. Accurate sequence files, complete metadata, appropriate annotation, correct taxonomy, careful validation, and good record keeping all contribute to a useful and reproducible database record.
- Understanding the GenBank submission process makes it easier to avoid common errors and obtain accession numbers efficiently. Once researchers understand the overall workflow, they can explore more specialized topics such as GenBank submission tools, preparing sequence files for GenBank, GenBank feature tables, WGS submissions, TSA submissions, and updating GenBank records in greater detail.