![]()
- BLASTN is a nucleotide sequence comparison program within the BLAST family of bioinformatics tools. The name BLASTN refers to a BLAST search in which a nucleotide query sequence is compared with nucleotide sequences in a selected database. It is widely used in molecular biology, genetics, genomics, microbiology, evolutionary biology, and bioinformatics to identify DNA sequences, find similar genes or genomic regions, validate experimental sequences, and investigate relationships between nucleotide sequences.
- BLASTN is particularly useful when a researcher has a DNA or RNA-derived nucleotide sequence and wants to determine whether similar sequences are already available in a sequence database. A query may come from Sanger sequencing, PCR, cloning, genome sequencing, transcriptome sequencing, environmental sequencing, or another molecular biology experiment. The sequence can then be compared against appropriate nucleotide databases to identify potentially significant matches.
- The basic principle of BLASTN is local nucleotide sequence similarity searching. Instead of requiring two sequences to match from beginning to end, BLASTN searches for regions where the query and database sequences are sufficiently similar. This is important because biological sequences can differ in length and may contain substitutions, insertions, deletions, conserved regions, or other differences while still sharing biologically meaningful sequence similarity.
- In a BLASTN search, the sequence submitted by the researcher is called the query sequence. The sequences in the database are the subject sequences. BLASTN compares the query against the selected nucleotide database and reports subject sequences containing regions that resemble the query. The resulting matches are displayed as BLAST hits and can be examined through their alignments and statistical information.
- BLASTN is commonly accessed through the NCBI BLAST web interface, where researchers can enter or upload a nucleotide sequence and select an appropriate nucleotide database. NCBI also provides BLAST+ software for researchers who need to perform searches locally or incorporate BLASTN into automated bioinformatics workflows.
- One of the simplest uses of BLASTN is identifying an unknown DNA sequence. For example, a researcher may sequence a DNA fragment obtained from a PCR experiment but not know exactly which region was amplified. By submitting the sequence to BLASTN, the researcher can search for similar nucleotide sequences in public databases. Strong and appropriately evaluated matches can provide evidence about the identity of the fragment.
- BLASTN can also be used to determine whether a sequence corresponds to a particular gene. If a nucleotide sequence produces strong matches to a known gene in several organisms, the researcher can examine the matching records and alignments to determine whether the unknown sequence is likely to represent that gene or a related sequence.
- The connection between BLASTN and GenBank is especially important. GenBank is a major public repository of nucleotide sequence information, whereas BLASTN is a tool for comparing nucleotide sequences. A BLASTN search can identify similar sequences, and the researcher can then examine the corresponding GenBank records to investigate the organism, accession number, sequence description, annotated genes, coding sequences, RNA features, and other information associated with the matching sequences.
- A typical BLASTN workflow begins by obtaining a nucleotide sequence of sufficient quality. The sequence is usually represented using the standard nucleotide alphabet and may be provided in FASTA format. The researcher then chooses a BLASTN search appropriate for the biological question, selects a suitable nucleotide database, submits the query, and evaluates the resulting hits.
- The choice of nucleotide database can have a substantial effect on the results. A broad nucleotide database may be appropriate when the identity of an unknown sequence is not known. A more specialized database or taxonomic restriction may be preferable when the researcher is investigating a particular group of organisms or a specific biological problem.
- The choice of BLASTN search task is also important. NCBI provides different nucleotide search strategies that are optimized for different situations. For example, megablast is designed for finding very similar nucleotide sequences efficiently and is particularly useful for comparisons involving closely related sequences. Other BLASTN strategies can provide greater sensitivity when the sequences are more divergent.
- A researcher should therefore consider the expected degree of similarity before selecting a search strategy. If the query is expected to be nearly identical to a reference sequence, a highly similar-sequence search can be appropriate. If the query may be more distantly related to available sequences, a more sensitive search strategy may be preferable.
- The length of the query sequence also matters. Long nucleotide sequences often provide more information for identifying a sequence than very short fragments. Short sequences can match multiple database sequences or may not contain enough variable positions to distinguish closely related organisms. NCBI provides search options intended for relatively short nucleotide queries, but the results still require careful biological interpretation.
- Once the BLASTN search has finished, the results generally contain a collection of sequence hits. Each hit represents a database sequence with a region similar to some part of the query. The most useful hits are not necessarily determined by a single number. Researchers should examine several measures together.
- One important measure is percentage identity. This describes the proportion of aligned nucleotide positions that are identical between the query and subject sequence. A high percentage identity indicates strong similarity within the aligned region, but percentage identity alone is not sufficient to determine whether a sequence is a good match.
- For example, imagine that one BLASTN hit has 100% identity but aligns with only 25 nucleotides of a 1,000-nucleotide query. Another hit has 98% identity across 980 nucleotides. The second result may provide much stronger evidence about the identity of the overall sequence because the similarity extends across nearly the entire query.
- This illustrates the importance of query coverage. Query coverage indicates how much of the query sequence is included in the alignment. High identity combined with high query coverage is generally much more informative than high identity over a very short region.
- The E-value, or expect value, is another important statistic reported by BLAST. It estimates how many matches with a score at least as good as the observed match would be expected to occur by chance under the search conditions. Smaller E-values generally indicate stronger statistical evidence for the observed similarity.
- The bit score provides another measure of alignment quality. Higher bit scores generally correspond to stronger alignments, and because the bit score is normalized, it is useful when comparing results obtained under different search conditions. E-value and bit score should nevertheless be interpreted alongside identity, coverage, alignment length, and biological context.
- The actual alignment is one of the most useful parts of a BLASTN result. It shows how the query sequence corresponds to the subject sequence and allows the researcher to examine matching nucleotides, substitutions, insertions, deletions, and the coordinates of the aligned region.
- Examining the alignment can reveal whether the similarity covers the expected biological region. For example, if the researcher expects a complete gene but the BLASTN alignment covers only a small conserved portion, the result may not be sufficient to establish that the entire query represents that gene.
- BLASTN results can also provide information about the orientation of the alignment. A nucleotide sequence may align to the same strand or to the reverse-complementary strand of a database sequence. This is important when interpreting genomic coordinates and sequence orientation.
- After identifying promising BLASTN hits, researchers should examine the corresponding database records. A GenBank record can provide valuable context that is not apparent from the BLAST alignment alone. The record may identify the organism, describe the sequence, provide an accession number, and contain annotated features such as genes and CDS regions.
- For example, if a query sequence produces a strong match to a bacterial gene, opening the corresponding GenBank record may reveal the gene name, coding sequence coordinates, organism, strain, and other annotations. This information can help determine whether the BLASTN result is biologically plausible.
- BLASTN is widely used for sequence identification. In microbiology, for example, researchers may compare a nucleotide sequence from an unknown organism against known microbial sequences. In biodiversity studies, DNA barcode sequences can be compared with reference sequences to investigate the possible identity of specimens. In molecular biology, cloned or amplified DNA fragments can be compared against reference sequences to confirm their expected identity.
- However, BLASTN should not be regarded as an automatic species-identification system. Closely related organisms may have highly similar nucleotide sequences, while the selected gene may not contain enough variation to distinguish them. Reliable identification may require additional genes, complete genomes, phylogenetic analysis, taxonomic information, or other independent evidence.
- BLASTN can also support gene annotation. When a newly assembled genomic sequence contains an unannotated region, similarity searches can help identify regions resembling previously characterized genes. Such evidence can contribute to annotation pipelines, although automated or similarity-based annotation should be reviewed appropriately before biological conclusions are drawn.
- Another application is sequence validation. Researchers can compare experimentally determined nucleotide sequences with expected reference sequences using BLASTN. A strong full-length match may support the expected identity of the sequence, whereas unexpected differences or unexpected database matches may indicate the need for further investigation.
- BLASTN is also useful for comparing sequences between organisms. Researchers can identify conserved regions, sequence variations, insertions, deletions, and other differences. These comparisons can contribute to studies of gene conservation, molecular evolution, population variation, and comparative genomics.
- In some cases, BLASTN can help identify potential contamination or unexpected sequence sources. For example, a sequence obtained from an experiment may produce strong matches to an organism that was not expected to be present. Such a finding can prompt investigation of laboratory contamination, sample handling, sequencing artifacts, or other explanations. A BLAST result alone, however, should not be considered definitive proof of contamination.
- BLASTN can also be used in conjunction with multiple sequence alignment and phylogenetic analysis. BLASTN can help researchers find candidate homologous sequences, which can then be retrieved from GenBank or other databases and subjected to more detailed comparative analyses. The BLAST search is therefore often an initial step rather than the final stage of an evolutionary study.
- One important distinction is that BLASTN identifies sequence similarity, not necessarily identical biological function. Two nucleotide sequences may be similar because they share evolutionary ancestry, occur in conserved genomic regions, or contain a common functional element. Conversely, a highly similar sequence may have a different biological role in a different genomic context. Interpretation therefore requires more than simply selecting the first result in the BLAST output.
- Database quality is another consideration. Public nucleotide databases contain enormous numbers of sequences from many organisms and projects, but sequences and annotations can differ in completeness and reliability. Researchers should consider the source and annotation of important reference sequences and, when possible, compare multiple independent hits.
- The taxonomic distribution of BLASTN hits can also provide useful information. If a query sequence matches highly similar sequences from several closely related organisms, the researcher may be able to identify the likely taxonomic group. If equally strong matches occur across very different organisms, the sequence may represent a highly conserved gene or another widely distributed sequence.
- Researchers should also be cautious when interpreting BLASTN results from repetitive or low-complexity sequences. Such regions may produce numerous matches that are not particularly informative for identifying the biological origin of the entire sequence. Examining the alignment and the biological location of the matching region is therefore important.
- For larger datasets, BLASTN can be run using BLAST+, NCBI’s command-line BLAST software. Local BLASTN searches are useful when researchers need to process many sequences, repeatedly search against a particular database, or integrate sequence comparison into a computational pipeline.
- BLAST+ also makes it possible to construct and search custom nucleotide databases. This can be useful when a researcher has a specific collection of reference sequences, such as sequences from a particular organism group, project, laboratory collection, or genome dataset.
- A conceptual BLASTN workflow can therefore be summarized as:
- Nucleotide sequence → choose BLASTN → select database and search strategy → perform search → examine hits → evaluate identity and coverage → inspect alignment → open relevant sequence records → interpret biological significance.
- Consider a researcher who obtains an 800-base-pair sequence from an experiment. The sequence is submitted to BLASTN, and several hits are returned. One hit has 99% identity over 790 bases and a very low E-value, while another has 100% identity over only 40 bases. The first hit would generally deserve much more attention because the similarity extends across nearly the entire query. The researcher would then inspect the associated database record to determine what gene or genomic region the sequence represents.
- This example demonstrates why BLASTN results must be interpreted as a combination of alignment quality and biological context. A high percentage identity is useful, but it becomes much more informative when accompanied by high query coverage, an appropriate alignment length, a strong statistical score, and a biologically credible reference sequence.
- BLASTN is therefore not simply a method for finding the “closest sequence.” It is a tool that helps researchers explore the relationships between nucleotide sequences. The final interpretation depends on the research question, the sequence being studied, the database searched, the search strategy, and the quality of the available reference information.
- BLASTN has become an essential tool because nucleotide sequence data are now generated on an enormous scale. From a short PCR product to an assembled genome, researchers frequently need to determine whether a sequence resembles something already known. BLASTN provides a practical way to make that initial connection.
- For beginners, the most important concepts to understand are the query sequence, subject sequence, nucleotide database, BLASTN search task, alignment, percentage identity, query coverage, E-value, and bit score. Understanding these concepts makes it much easier to evaluate whether a BLASTN result is genuinely informative.
- In summary, BLASTN is a nucleotide sequence similarity-searching tool used to compare DNA or RNA-derived nucleotide sequences with nucleotide sequences in a database. It is widely used for sequence identification, gene analysis, sequence validation, comparative genomics, annotation, and many other biological applications. When combined with GenBank, BLASTN allows researchers to move from a sequence similarity result to detailed information about the corresponding biological sequence and its annotations.
- BLASTN is only one part of the broader BLAST family. The next level of understanding involves learning how to choose BLASTN search settings, interpret E-values and bit scores, evaluate query coverage and percentage identity, distinguish BLASTN from megablast, and interpret complete BLAST results. These topics can be explored in dedicated articles as part of the wider BLAST and GenBank sequence-analysis series.