![]()
- GenBank has grown from a relatively small collection of nucleotide sequences into one of the world’s largest public repositories of biological sequence information. Its release history provides a useful way to understand the development of modern sequence databases, the expansion of DNA sequencing, and the increasing importance of genomic data in biology and bioinformatics. The database is maintained by the National Center for Biotechnology Information (NCBI) and forms one part of the International Nucleotide Sequence Database Collaboration (INSDC), together with the DNA Data Bank of Japan (DDBJ) and the European Nucleotide Archive (ENA). GenBank releases provide periodic snapshots of the database and document changes in sequence data, records, formats, and database infrastructure. NCBI currently issues a GenBank release approximately every two months.
- The history of GenBank begins in the early 1980s, when nucleotide sequence databases were still relatively small and the amount of experimentally determined DNA sequence was limited. GenBank Release 3, issued in December 1982, is the starting point used in NCBI’s historical growth statistics. It contained only 606 sequence records comprising 680,338 bases. Although tiny by modern standards, this early collection represented an important step toward creating a centralized resource in which nucleotide sequences could be stored, shared, and reused by researchers. NCBI’s historical statistics show how dramatically the database expanded from this starting point.
- During the first several years, GenBank expanded rapidly as researchers increasingly determined and published DNA sequences. By Release 14 in November 1983, the database contained 2,427 sequences and more than 2.27 million bases. Release 20 in May 1984 contained approximately 3.0 million bases, while Release 40 in February 1986 had grown to more than 5.9 million bases. By Release 48 in February 1987, the database contained more than 10.9 million bases. These early milestones illustrate how quickly publicly available nucleotide sequence information was beginning to accumulate even before the era of large-scale genome sequencing.
- An important part of GenBank’s early development was its increasing integration with international sequence databases. GenBank did not develop as an isolated American resource. It became part of an international framework involving the European sequence database and the DNA Data Bank of Japan, eventually forming the coordinated system now known as the International Nucleotide Sequence Database Collaboration. This international exchange helped prevent major duplication of effort and established a mechanism through which sequence data submitted to one participating database could become available through the others.
- The 1990s marked a major period of expansion. GenBank Release 100, issued on April 15, 1997, contained 1,274,747 reported sequences and approximately 842.9 million bases. The release documentation also illustrates how GenBank was already incorporating sequence information from multiple sources, including direct submissions, literature-derived records, international collaborators, and other sequence resources.
- The growth continued rapidly toward the end of the 1990s and into the early 2000s. Release 115 in December 1999 contained more than 4.65 billion bases and over 5.35 million sequences. Release 121 in December 2000 had already exceeded 11.1 billion bases and 10 million sequences. By Release 127 in December 2001, the database contained approximately 15.85 billion bases and nearly 15 million sequences. This period coincided with the rapid expansion of molecular biology databases, high-throughput sequencing, genome projects, and computational biology.
- The early 2000s were particularly important because genome sequencing projects began contributing increasingly large quantities of data. Instead of receiving only individual genes or relatively short sequences, public repositories increasingly had to accommodate large genomic datasets and new forms of sequence data. This changing scale eventually contributed to the development and use of specialized data categories such as Whole Genome Shotgun (WGS) records. Understanding these categories is important when interpreting historical GenBank statistics because traditional sequence records and large-scale set-based datasets are not always counted in the same way.
- By the middle of the 2000s, GenBank had entered a new phase of large-scale growth. Release 150, issued in October 2005, contained approximately 53.7 billion bases and 49.2 million reported sequences in the traditional GenBank statistics. By Release 160 in June 2007, the database contained approximately 97.1 billion bases and more than 73 million sequences. Release 168 in October 2008 had reached approximately 97.4 billion bases and more than 96 million sequences, illustrating the enormous acceleration in sequence-data production during the genomics era.
- Release 200, issued on February 15, 2014, represents another useful historical milestone. At that point, the traditional GenBank statistics reported approximately 157.9 billion bases across more than 171 million loci. By then, GenBank had become a fundamental infrastructure for molecular biology, genomics, sequence analysis, phylogenetics, genome annotation, and many other areas of biological research.
- The growth of GenBank cannot be understood simply as a steady increase in the number of individual genes. Modern sequencing technologies changed both the volume and nature of the data entering public repositories. High-throughput sequencing enabled researchers to generate complete genomes, draft genomes, transcriptome assemblies, metagenomic datasets, and other large collections of nucleotide sequences. Consequently, later GenBank releases increasingly reflected not only the growth of traditional records but also the enormous expansion of WGS, Transcriptome Shotgun Assembly (TSA), and other large-scale sequence datasets.
- This distinction is important when examining GenBank statistics. NCBI’s historical statistics specifically distinguish traditional GenBank records from WGS data, while more recent releases also report additional components such as TSA and TLS. Therefore, a comparison of GenBank database size across different historical periods should always consider which categories of sequence data are being counted. A database containing billions of sequences does not necessarily contain billions of independent biological discoveries, because sequence records can represent related sequences, assemblies, different versions, or components of larger datasets.
- Another important milestone in GenBank’s development was the increasing standardization of sequence annotation and record structure. As the database grew, it became increasingly important to maintain consistent representations of genes, coding sequences, RNA features, source information, taxonomy, publications, and other biological annotations. The GenBank flat file format became an important mechanism for distributing and exchanging annotated sequence records, while standardized feature definitions supported interoperability between international sequence databases.
- GenBank releases also became important historical records of database development. Each release provides information about changes to the database, new sequence data, updated records, organizational changes, and other technical developments. NCBI maintains an extensive archive of release notes, allowing researchers to examine the evolution of GenBank across many years. The release archive shows the regular progression of releases, including Release 100 in 1997, Release 150 in 2005, Release 200 in 2014, Release 250 in 2022, and subsequent releases through the present period.
- The transition from the 2000s to the 2010s and 2020s illustrates how sequencing technology transformed the scale of public sequence repositories. The emergence of next-generation sequencing and increasingly efficient genome assembly methods produced sequence datasets at a scale that would have been difficult to imagine when GenBank Release 3 contained only 680,338 bases. The resulting growth has affected not only storage requirements but also sequence searching, annotation, data retrieval, computational analysis, and the design of bioinformatics workflows.
- Recent GenBank releases demonstrate just how dramatic this transformation has become. GenBank Release 273.0, dated August 15, 2026, contained approximately 60.07 trillion bases and 6.68 billion records across its reported components. The release included approximately 267.4 million traditional records containing 8.24 trillion base pairs, more than 5.13 billion WGS records containing approximately 50.83 trillion base pairs, about 1.08 billion TSA records containing approximately 923.9 billion base pairs, and approximately 193.6 million TLS records containing about 80.2 billion base pairs.
- The comparison between recent releases also shows that GenBank continues to grow extremely rapidly. Between Releases 272.0 and 273.0, the traditional component increased by more than 618 billion base pairs and approximately 3.17 million sequence records. During the same interval, the WGS component increased by approximately 1.75 trillion base pairs and more than 138 million records. TSA increased by approximately 12.5 billion base pairs and more than 14 million records, while TLS increased by more than 608 million base pairs and approximately 1.18 million records.
- NCBI’s historical statistics summarize this extraordinary long-term expansion by noting that, from 1982 to the present, the number of bases in traditional GenBank has approximately doubled every 18 months. The precise growth rate has varied across different periods, but the overall trend demonstrates the exponential expansion of publicly available nucleotide sequence information.
- The changing scale of GenBank has also changed the way scientists use the database. In the early years, researchers could work with relatively small collections of nucleotide sequences. Today, GenBank is routinely used for BLAST searches, gene identification, sequence comparison, genome annotation, phylogenetic analysis, DNA barcoding, microbial identification, comparative genomics, biodiversity research, and many other applications. The enormous growth of the database has increased the probability that a newly analyzed sequence will have related sequences available for comparison, while simultaneously making careful database searching and result interpretation more important.
- The history of GenBank is therefore closely connected with the history of modern bioinformatics. As sequence datasets expanded, researchers developed increasingly sophisticated methods for searching, retrieving, annotating, comparing, and analyzing nucleotide sequences. Tools such as NCBI Entrez, BLAST, sequence retrieval systems, genome browsers, and programmatic interfaces have become essential because manually examining individual records is no longer practical at modern database scale.
- Another important consequence of database growth is the increasing importance of accession numbers and sequence versions. As records are updated, researchers need persistent identifiers to locate particular sequences and to distinguish versions when necessary. For this reason, understanding GenBank accession numbers and accession versions is an important part of working with historical and current GenBank data.
- GenBank’s release history also demonstrates why reproducibility matters in bioinformatics. A search performed against a rapidly changing sequence database may produce different results at different times. New sequences may be added, existing records may be updated, annotations may change, and database structures may evolve. Researchers therefore benefit from recording accession numbers, accession versions where appropriate, database resources, search methods, and relevant release information when documenting computational analyses.
- It is also important to distinguish GenBank from NCBI’s RefSeq database. GenBank is primarily an archival repository of publicly submitted nucleotide sequence data, whereas RefSeq is a curated reference sequence collection designed to provide nonredundant reference sequences. The two resources serve complementary purposes and should not be treated as interchangeable when interpreting database growth or sequence-analysis results.
- The historical progression from hundreds of thousands of bases in the early 1980s to tens of trillions of bases today reflects several major developments in biology: the expansion of molecular sequencing, international data sharing, genome projects, high-throughput sequencing, improved computational infrastructure, and the increasing use of sequence information across biological disciplines. GenBank’s release history therefore represents much more than a record of database growth; it is also a record of the transformation of biological research from relatively small-scale sequence analysis into data-intensive genomics.
- For students and researchers, understanding this history provides useful context for interpreting modern GenBank records and statistics. A current database containing billions of records is the result of more than four decades of accumulated sequence submissions, updates, international data exchange, large-scale sequencing projects, and technological change. The historical release numbers provide a convenient framework for following this transformation from GenBank Release 3 in 1982 to the modern releases of the 2020s.
- GenBank continues to evolve as sequencing technologies, annotation methods, computational infrastructure, and biological research requirements change. Release 273.0 in August 2026 demonstrates that the database is still expanding at an extraordinary rate, with WGS and other large-scale datasets contributing substantially to overall growth. At the same time, NCBI continues to document changes through release notes and related technical information, making the release history an important resource for anyone interested in the development and use of public nucleotide sequence databases.
- In summary, the GenBank release history provides a timeline of the growth of public nucleotide sequence data from a small collection of hundreds of records in 1982 to a massive international resource containing trillions of nucleotide bases and billions of records. Major milestones such as Releases 3, 50, 100, 150, 200, 250, and the current 273 series illustrate successive stages in the development of sequence databases. Studying these milestones helps explain why GenBank has become central to modern genomics, molecular biology, evolutionary research, microbiology, biodiversity studies, and bioinformatics.