![]()
- UniProt, short for the Universal Protein Resource, is one of the most widely used resources for accessing information about proteins. It provides researchers, students, bioinformaticians, and life-science professionals with protein sequences together with information about their functions, structures, biological roles, molecular interactions, subcellular locations, and other characteristics. UniProt is particularly important in protein bioinformatics, where large amounts of biological sequence and functional information need to be organized and interpreted efficiently. The UniProt Knowledgebase (UniProtKB) is the central component of the resource and is designed to provide accurate, consistent, and richly annotated protein information.
- The history of UniProt is closely connected with the development of major protein sequence databases. UniProtKB incorporates the work of the Swiss-Prot and TrEMBL databases, bringing manually curated and computationally annotated protein records together within a common resource. Swiss-Prot, established in 1986, became the reviewed component of UniProtKB, while TrEMBL was introduced in 1996 to handle the rapidly increasing number of protein sequences generated by genome projects. This combination allows UniProt to provide both highly curated information and large-scale computationally generated data.
- The central resource within UniProt is UniProtKB, the UniProt Knowledgebase. Each UniProtKB entry can contain the protein sequence, protein name, organism information, references, functional descriptions, sequence features, and links to other biological databases. These entries can also contain information about protein domains, active sites, binding sites, post-translational modifications, transmembrane regions, signal peptides, and subcellular locations. The richness of these annotations makes UniProt useful not simply as a sequence database but as a comprehensive source of protein knowledge.
- An important aspect of UniProtKB is the distinction between UniProtKB/Swiss-Prot and UniProtKB/TrEMBL. Swiss-Prot contains reviewed records that have undergone manual annotation and evaluation by expert curators. TrEMBL contains unreviewed records that have primarily been processed through computational annotation. The distinction is important when evaluating the level of evidence and curation associated with a protein record. Computationally generated annotations are valuable for handling the enormous volume of sequence data, while manually reviewed records provide a higher level of curated biological interpretation.
- Protein sequence data forms the foundation of UniProt. A large majority of UniProtKB protein sequences originate from translations of coding sequences submitted to the International Nucleotide Sequence Database Collaboration, which includes ENA, GenBank, and DDBJ. UniProt integrates these sequences and associates them with additional biological information. Other sources can also contribute sequences, including experimentally determined protein sequences and information derived from scientific literature and other databases.
- Another major aspect is protein annotation. Annotation adds biological meaning to a protein sequence by describing what the protein is believed to do, where it is located, what processes it participates in, and what molecular characteristics it possesses. UniProt annotations can include protein function, catalytic activity, biological pathways, subcellular location, domains, sequence features, cofactors, interactions, and other information. Annotation may be supported by experimental evidence, computational analysis, similarity-based inference, or other forms of evidence, and UniProt provides indications of annotation quality and evidence.
- Protein function is one of the most important types of information available through UniProt. A UniProt entry may describe the biological role of a protein, its molecular activity, the reactions it catalyzes, or the cellular processes in which it participates. Functional information can help researchers move from a raw amino acid sequence toward an understanding of the biological role of the encoded protein. However, the evidence supporting a functional annotation should always be considered when interpreting the information.
- UniProt also provides extensive information about protein names and identifiers. A protein can have a recommended name, alternative names, gene names, and unique UniProt identifiers. These identifiers are particularly useful when connecting information across publications, databases, computational pipelines, and research projects. Stable identifiers make it easier to retrieve the same protein record even when descriptive names vary between resources.
- Sequence features provide another important layer of information. UniProt can identify or describe regions of a protein sequence associated with biological or structural characteristics, including signal peptides, transmembrane regions, active sites, binding sites, domains, repeats, and other regions of interest. These features allow users to examine a protein not simply as a continuous amino acid sequence but as a collection of biologically meaningful regions.
- Protein domains and families are also closely connected with UniProt annotation. Domains are regions of proteins that can have characteristic structures or functions and may occur across different proteins and organisms. UniProt links protein records with resources such as InterPro, Pfam, PROSITE, PANTHER, Gene3D, and other classification resources. These cross-references help researchers investigate evolutionary relationships, conserved regions, protein families, and functional characteristics.
- Gene Ontology (GO) is another important component associated with UniProt. Gene Ontology provides standardized terminology for describing molecular functions, biological processes, and cellular components. UniProt entries can be connected to GO annotations, allowing researchers to use a common vocabulary when comparing proteins and interpreting biological datasets. GO information is particularly useful in functional genomics, enrichment analysis, and large-scale biological studies.
- Protein localization describes where a protein is found within a cell or organism. UniProt may provide information about cellular compartments and subcellular locations, such as the nucleus, cytoplasm, mitochondria, membrane, or extracellular space. Understanding localization can provide important clues about protein function because proteins generally perform their biological roles in specific cellular environments.
- Post-translational modifications (PTMs) represent another important aspect of UniProt. Proteins can be chemically modified after translation, and these modifications can influence their activity, stability, localization, interactions, or degradation. UniProt records may contain information about modifications such as phosphorylation, acetylation, glycosylation, and other types of processing. Sequence features can indicate the positions at which experimentally characterized or predicted modifications occur.
- UniProt also contains information about protein isoforms and alternative forms of proteins. Different isoforms can arise through mechanisms such as alternative splicing and may differ in sequence, localization, regulation, or biological function. UniProt can provide separate information about isoforms when sufficient evidence is available. In the 2026_02 release, for example, UniProtKB contained more than 41,000 isoforms associated with reviewed records.
- Protein existence evidence is another useful feature for evaluating UniProt records. UniProt classifies protein existence according to the available evidence, ranging from evidence at the protein level to evidence at the transcript level, inference from homology, prediction, or uncertain existence. This distinction is important because the presence of a protein sequence in a database does not necessarily mean that the protein has been experimentally observed.
- Taxonomy allows UniProt users to explore proteins according to their biological origin. Protein records are associated with organisms and taxonomic classifications, making it possible to study proteins from bacteria, archaea, viruses, fungi, plants, animals, and other organisms. Taxonomic information is particularly useful when comparing homologous proteins across species or restricting a search to a particular organism.
- The UniProt Proteomes resource provides another way to explore protein information. A proteome represents the set of proteins associated with an organism, generally based on a completely sequenced genome. UniProt provides proteome sets that can contain both reviewed and unreviewed protein entries, depending on the organism and the level of available curation. Proteome data is useful for studying the complete protein complement of an organism and for comparative genomics.
- Beyond UniProtKB and Proteomes, UniProt is associated with resources that help manage protein sequence redundancy and sequence history. UniRef, or UniProt Reference Clusters, groups similar protein sequences to reduce redundancy and make large-scale sequence analysis more efficient. UniParc, the UniProt Archive, provides a comprehensive sequence archive that helps maintain a historical record of protein sequences and their database sources. Together, these resources complement the functional information available in UniProtKB.
- Cross-references to external databases are a major strength of UniProt. A UniProt entry can link to resources covering protein structures, domains, pathways, taxonomy, genomes, gene ontology, protein families, literature, and other biological information. Important connections include databases and resources such as PDB, InterPro, Pfam, GO, PANTHER, PROSITE, AlphaFoldDB, and many others. These links allow UniProt to function as a central starting point from which researchers can explore different aspects of a protein.
- Protein structure information can also be explored through UniProt’s connections to structural resources. Researchers can use a UniProt protein record to move toward experimentally determined structures or computationally predicted structures. Connections with resources such as the Protein Data Bank and AlphaFoldDB can therefore help bridge the gap between protein sequence and three-dimensional structure.
- Literature and scientific references are another important component of UniProt. Protein entries can contain references to scientific publications that provide experimental evidence or other information used during annotation. Reviewing these references can help researchers distinguish experimentally demonstrated biological properties from predictions or computational inferences. Literature links also provide a pathway from a database record to the original scientific evidence.
- UniProt provides several search and retrieval tools that allow users to find proteins using identifiers, gene names, protein names, organisms, sequences, functional terms, and other criteria. Researchers can search individual records or construct more complex queries to retrieve groups of proteins. Search functionality is particularly valuable when working with large datasets because users can combine biological and taxonomic criteria to narrow down results.
- For researchers working with large datasets, UniProt downloads and APIs provide programmatic access to protein information. Data can be retrieved in formats suitable for computational analysis, allowing UniProt information to be incorporated into bioinformatics pipelines, scripts, databases, and research workflows. This makes UniProt useful not only as a web-based reference resource but also as a source of machine-readable biological data.
- UniProt data quality and curation are central to its usefulness. Manual curation, automated annotation, evidence classification, sequence analysis, literature integration, and regular database updates all contribute to the continuing development of the resource. Because the volume of biological sequence data is enormous, UniProt combines expert manual curation with computational approaches to make information available at scale while maintaining a distinction between reviewed and unreviewed records.
- The scale of UniProt illustrates why such a resource is necessary. In the UniProtKB 2026_02 release, there were approximately 149.8 million entries, including about 575,000 reviewed Swiss-Prot entries and more than 149 million unreviewed TrEMBL entries. The database represented more than 230,000 species. These figures demonstrate the enormous growth of protein sequence information and the challenge of organizing it into a resource that can be searched and interpreted by researchers.
- UniProt is widely used in bioinformatics, molecular biology, genomics, proteomics, structural biology, biotechnology, and biomedical research. Researchers may use it to identify a protein, investigate its function, compare sequences between organisms, examine conserved regions, study protein families, investigate disease-related proteins, design experiments, or prepare datasets for computational analysis. Its combination of sequence information, functional annotation, evidence, literature, and external database links makes it an important starting point for many biological investigations.
- Overall, UniProt can be viewed as much more than a protein sequence database. It is an interconnected knowledge resource covering protein sequences, annotation, function, domains, families, structures, modifications, localization, taxonomy, proteomes, evidence, literature, identifiers, and external database links. Understanding these different aspects is essential for using UniProt effectively and for interpreting protein information correctly. In the following articles, each of these aspects can be examined in greater detail, including how the information is organized, how to search for it, how to interpret the evidence, and how UniProt data can be used in practical bioinformatics research.