GenBank Sequence Submission Tools: Preparing and Validating Your Data

Loading

  • Submitting nucleotide sequences to GenBank is an important step in making biological sequence data publicly available, reusable, and discoverable by the scientific community. However, successful submission involves more than simply uploading a DNA sequence. Researchers must select the appropriate submission pathway, prepare sequence files, provide accurate biological source information, add appropriate annotation, check the data for errors, and respond to validation or review requests. The quality of the information submitted at this stage directly affects the usefulness and reliability of the resulting GenBank record.
  • NCBI provides several tools and submission pathways because different types of sequence data require different preparation and validation procedures. The current NCBI Submission Portal helps researchers determine the appropriate destination based on whether their data consist of assembled nucleotide sequences, complete or draft genomes, transcriptome assemblies, or unassembled sequencing reads. In general, Submission Portal-GenBank is used for assembled nucleotide sequences other than prokaryotic and eukaryotic genomes and transcriptomes, while genome assemblies, transcriptome assemblies, and raw high-throughput sequencing reads have dedicated submission pathways such as Genome, TSA, and SRA.
  • Choosing the correct submission tool should therefore be the first part of the preparation process. A researcher submitting a single gene, a group of related loci, a viral sequence, a ribosomal RNA sequence, or another assembled nucleotide sequence may use the appropriate GenBank workflow. A complete or draft prokaryotic or eukaryotic genome should normally be submitted through the Genome pathway, whereas computationally assembled transcript sequences derived from sequencing reads belong in the Transcriptome Shotgun Assembly pathway. Unassembled high-throughput sequencing reads should be deposited in the Sequence Read Archive rather than submitted as ordinary GenBank sequences.
  • The current Submission Portal-GenBank (SP-GenBank) provides workflows organized around major types of sequenced material, including prokaryotes, eukaryotes, viruses, and synthetic constructs. Some sequence types have specialized automatic annotation, while other submissions require the submitter to provide feature annotation. Before beginning a submission, researchers should have their sequence data, biological source information, publication information where applicable, sequencing technology information, and required annotations available.
  • The sequence itself should normally be prepared in FASTA format. A FASTA file contains a definition line beginning with the > character followed by a sequence identifier and the nucleotide sequence. Sequence identifiers should be short, simple, and free of spaces. A single SP-GenBank submission can currently contain one FASTA file with up to 3,000 sequences, although fewer sequences may be appropriate when the sequences are particularly long.
  • Before uploading a FASTA file, researchers should carefully inspect the nucleotide sequences. Sequence names should be unique and consistently associated with the corresponding biological samples and metadata. Unexpected characters, accidental line breaks, duplicated sequences, truncated sequences, incorrect orientations, and sample mix-ups can create problems during submission or compromise the resulting records. The nucleotide sequence should also correspond exactly to the sequence used in the analyses and manuscript, where applicable. Maintaining a master copy of the original sequence data and a separate submission-ready version is a useful practice for scientific data management.
  • Sequence quality is equally important. Researchers should confirm that the sequences are biologically plausible and that obvious sequencing or assembly problems have been investigated before submission. Depending on the type of study, this may include examining chromatograms, read coverage, assembly quality, ambiguous bases, unexpected stop codons, frame disruptions, contamination, or other abnormalities. GenBank is a public archival database, so submitting an unresolved error can propagate incorrect information into subsequent research, analyses, and publications.
  • The next major component is biological source information. GenBank submissions require information describing where the sequence originated. This commonly includes the organism name and may include additional qualifiers such as isolate, strain, specimen, collection date, geographic location, host, tissue, or other information appropriate to the study. NCBI checks organism names against its Taxonomy database, and unfamiliar organism names may require additional information or clarification.
  • Source metadata should be checked carefully before submission because it provides the biological context necessary for interpreting a sequence. A nucleotide sequence without reliable source information may have considerably less scientific value than an equivalent sequence accompanied by well-documented metadata. Researchers should therefore make sure that sample identifiers, organism names, collection information, and other relevant source attributes are consistent across laboratory records, sequence files, manuscripts, BioProject or BioSample records, and the GenBank submission.
  • Sequence annotation is another central part of GenBank submission preparation. Depending on the sequence type, annotation may identify genes, coding sequences (CDSs), regulatory regions, ribosomal RNA, transfer RNA, internal transcribed spacers, or other biologically meaningful features. Some specialized SP-GenBank workflows can automatically annotate particular sequence types, whereas other submissions require the submitter to provide feature annotation.
  • Researchers should pay particular attention to the relationship between sequence coordinates and annotations. A coding sequence should correspond to the correct nucleotide interval and reading frame, while gene and RNA features should be positioned correctly on the submitted sequence. Qualifiers such as gene names, product descriptions, isolate information, and other biological descriptors should also be accurate and consistent. Understanding the GenBank feature table and GenBank sequence annotation system is therefore valuable before beginning a complex submission.
  • For larger or more complicated datasets, specialized preparation tools may be more appropriate than entering information manually. NCBI identifies table2asn as a command-line program that automates the creation of sequence records for GenBank and is used primarily for annotated genomes and large batches of sequences. It can help transform sequence and annotation information into submission-ready data and is particularly useful when processing many records or complex annotations.
  • The choice between a web-based submission workflow and a preparation tool depends on the nature and scale of the dataset. A researcher submitting a relatively small collection of assembled sequences may find the Submission Portal convenient because it provides an interactive environment for entering or uploading metadata and reviewing the submission. Larger projects with extensive annotation or many records may benefit from automated preparation and validation workflows. NCBI’s current submission system also provides different portals for genome, TSA, SRA, and other data types, making it important to select the tool according to the biological nature of the dataset rather than simply choosing the most familiar interface.
  • Validation is an essential part of the GenBank submission process. After a submission is entered, NCBI performs automated validation and manual review. Automated checks can identify problems in the submitted information, while manual review allows NCBI staff to examine submissions that require additional biological or technical assessment. A submission may therefore pass some automated checks but still require clarification or correction before accession numbers are assigned or the record is released.
  • Common validation problems can arise from inconsistent metadata, invalid or incomplete feature annotations, incorrect organism names, problematic sequence identifiers, missing required information, or discrepancies between sequences and their annotations. Some problems are straightforward formatting errors, whereas others may indicate a deeper biological issue. Researchers should treat validation messages as an opportunity to improve the quality of the submission rather than simply as technical obstacles.
  • The current Submission Portal provides status information as a submission progresses. A submission may initially be marked as having an error, may enter processing before accession numbers are assigned, may proceed through final review after accession numbers have been assigned, and eventually reach a processed status when GenBank processing is complete and the records are ready for release. NCBI may contact submitters if additional information or corrections are required.
  • An important change for researchers following older GenBank tutorials is the transition from BankIt to Submission Portal-GenBank. NCBI states that SP-GenBank has expanded to accept most submission types that were historically handled through BankIt. As of 2026, BankIt remains relevant for aligned sequence submissions requiring feature propagation because that functionality is not yet available in SP-GenBank, but NCBI is moving toward broader use of the Submission Portal and expects BankIt to be discontinued later in 2026. Researchers should therefore verify current NCBI instructions rather than relying on older tutorials that describe BankIt as the general GenBank submission tool.
  • Another important consideration is the relationship between GenBank submissions and BioProject and BioSample records. These resources can provide project-level and sample-level context for sequence datasets and improve the organization and discoverability of related data. NCBI indicates that BioProject, BioSample, and SRA accessions are generally optional for GenBank submissions, although particular submission circumstances may require them. The Submission Portal can also support creation of BioProject and BioSample records during sequence-data submission in appropriate cases.
  • Researchers preparing data for publication should also consider when the sequences need to become public. Many journals require sequence data described in a manuscript to be deposited in a public nucleotide database. GenBank allows sequences submitted before publication to be kept confidential upon request, and NCBI states that accession numbers are typically assigned within two working days, although the complete processing time can vary. Accession numbers can then be included in the manuscript so readers can retrieve the underlying sequence data.
  • Privacy and responsible data submission are particularly important when sequence information is associated with humans. Submitters should ensure that information included with a GenBank record does not reveal the personal identity of a source. NCBI also states that, effective June 16, 2026, GenBank no longer accepts personal sequence data from private individuals, including personal human chromosome, genome, or mitochondrial DNA data in the examples specified by NCBI. Researchers working with human-associated sequence data should therefore review the current NCBI requirements before submission.
  • A useful preparation strategy is to treat a GenBank submission as a structured data-validation project rather than as a simple file-upload task. First determine which NCBI repository and submission pathway is appropriate. Then prepare the sequence files, verify sequence quality and identifiers, assemble accurate source metadata, prepare annotation when required, and check the relationship between sequences and biological samples. After the submission is uploaded, carefully review validation messages and make corrections before the records proceed through final processing.
  • Keeping a complete local copy of the submitted FASTA files, metadata tables, annotation files, validation reports, correspondence with NCBI, and final accession numbers is also good scientific practice. These materials provide a reproducible record of what was submitted and make it easier to update GenBank records later if an error is discovered. Once accession numbers have been assigned, researchers should use the appropriate GenBank record update procedure rather than creating an unnecessary duplicate submission.
  • The most effective GenBank submission workflow can therefore be summarized as a sequence of connected stages: choose the correct submission pathway, prepare the sequences, verify sequence quality, prepare biological source metadata, create or verify annotation, validate the submission, correct detected problems, complete NCBI review, obtain accession numbers, and retain the final submission documentation. Each stage contributes to the accuracy and long-term usefulness of the resulting GenBank records.
  • Understanding GenBank submission tools is especially important because the tools and submission pathways continue to evolve. Current NCBI guidance should always take precedence over older instructions, particularly when tutorials refer to legacy tools such as BankIt or older genome-submission procedures. By combining careful sequence preparation, accurate metadata, appropriate annotation, and systematic validation, researchers can substantially reduce submission errors and create high-quality sequence records that remain useful to the scientific community for years to come.
Author: admin

Leave a Reply

Your email address will not be published. Required fields are marked *