![]()
- UniProtKB, or the UniProt Knowledgebase, is the central component of UniProt and one of the most important resources for accessing information about proteins. It brings together protein sequences and a wide range of biological and functional information, allowing researchers to investigate what proteins are, what they do, where they are located, how they are regulated, and how they relate to other proteins. UniProt describes UniProtKB as a central hub for functional protein information, with an emphasis on accurate, consistent, and richly annotated data.
- The primary purpose of UniProtKB is to transform protein sequence information into biologically useful knowledge. A protein sequence by itself consists mainly of a string of amino acids, but researchers generally need much more information to understand its biological significance. UniProtKB therefore combines the sequence with information such as the protein name, gene name, organism, function, cellular location, biological process, molecular activity, sequence features, scientific references, evidence, and links to other biological databases.
- Every UniProtKB entry represents a protein record containing a collection of information associated with a particular protein sequence. Core information includes the amino acid sequence, protein name or description, taxonomic information, and citation information. Additional annotations can describe many other characteristics of the protein. These annotations are organized so that researchers can move from basic identification of a protein toward more detailed biological interpretation.
- A particularly important characteristic of UniProtKB is that it contains two major sections: UniProtKB/Swiss-Prot and UniProtKB/TrEMBL. Swiss-Prot contains reviewed records that have been manually annotated, while TrEMBL contains unreviewed records that are primarily computationally annotated. This distinction allows UniProtKB to accommodate the enormous quantity of available protein sequences while maintaining a separate collection of records that have undergone extensive expert review.
- UniProtKB/Swiss-Prot is the reviewed section of the Knowledgebase. Its records are manually annotated by expert curators who evaluate experimental findings, scientific literature, computational analyses, and other available evidence. Manual annotation involves critical evaluation rather than simply copying information from another database. Curators review the available evidence and organize it into a standardized protein record.
- UniProtKB/TrEMBL is the unreviewed section of UniProtKB. It contains a very large collection of protein sequences that have been computationally analyzed and annotated but have not necessarily undergone complete manual review. This approach makes it possible to provide information about a vast number of proteins while allowing selected records to receive additional manual curation when appropriate.
- The difference between reviewed and unreviewed records is important when interpreting UniProt information. A reviewed record does not simply mean that every individual statement has been experimentally demonstrated, but it indicates that the entry has undergone expert manual annotation. An unreviewed record may contain valuable computationally generated information, but users should pay particular attention to the evidence associated with individual annotations before treating them as experimentally established facts.
- The protein sequence is one of the fundamental components of a UniProtKB entry. More than 95% of the protein sequences provided by UniProtKB are derived from translations of coding sequences submitted to the ENA, GenBank, and DDBJ resources of the International Nucleotide Sequence Database Collaboration. These translated sequences are automatically incorporated into the unreviewed TrEMBL section, where they can subsequently be considered for further annotation and curation.
- UniProtKB can contain information from several additional sources besides translated nucleotide sequences. These include sequences associated with the Protein Data Bank, experimentally determined protein sequences, sequences obtained from scientific literature, and sequences derived from gene prediction that have not been submitted to the major nucleotide sequence repositories. This integration helps UniProtKB provide a broad representation of known protein sequence information.
- Protein annotation is what gives biological meaning to the sequence stored in a UniProtKB entry. Annotation can describe protein function, catalytic activity, biological processes, subcellular location, interactions, domains, post-translational modifications, sequence variants, and many other characteristics. UniProt aims to standardize this information where possible by using controlled vocabularies, ontologies, classifications, and links to specialized databases.
- Functional annotation is one of the most valuable aspects of UniProtKB. It can describe the biological role of a protein and, for enzymes, may include information about catalytic activity and the reactions that the protein performs. Other functional information can describe molecular interactions, biological pathways, disease associations, biotechnology applications, or other biologically relevant properties.
- Evidence information helps users understand how individual annotations have been supported. UniProt distinguishes between experimental evidence, computational evidence, and other forms of evidence. This is particularly important because not every annotation has the same level of experimental support. Understanding evidence attribution is therefore essential when using UniProtKB for research, database construction, or computational analysis.
- Manual curation is a major feature that distinguishes the reviewed section of UniProtKB. During manual annotation, expert biologists critically evaluate experimental and predicted information, extract relevant information from scientific literature, verify computational results, integrate large-scale datasets, and update records as new evidence becomes available. This process helps produce consistent and biologically meaningful protein records.
- UniProt also uses automatic annotation to handle the enormous volume of protein sequences that cannot all be manually reviewed. Automated systems include UniRule, which consists of manually curated annotation rules that can propagate annotations under defined conditions, and ARBA, an Association-Rule-Based Annotator that provides automatic classification and annotation. These approaches allow computational annotation to operate at a much larger scale than manual curation alone.
- A UniProtKB entry contains several types of protein identifiers. The UniProt accession number is a stable identifier used to refer to a particular protein record, while the entry name provides another form of identification. Gene names, protein names, organism information, and other identifiers help researchers connect the UniProtKB record with information from publications and other biological databases.
- Protein names and gene names are important for identifying proteins across different biological resources. A protein can have a recommended name as well as alternative names or synonyms. Gene names provide an additional way to search for and identify records. Because naming conventions can vary between organisms and scientific publications, UniProtKB brings several types of identifiers together within the same entry.
- Taxonomic information connects each protein to the organism from which it originates. This makes it possible to search and filter protein records by species, taxonomic group, or other organism-related criteria. Taxonomy is particularly useful when comparing proteins from different organisms or when a research project is limited to a particular species or lineage.
- Sequence features provide information about specific regions or positions within a protein sequence. Examples include active sites, binding sites, signal peptides, transmembrane regions, domains, modified residues, disulfide bonds, sequence variants, and other biologically relevant features. These annotations help researchers interpret the sequence at a more detailed level rather than treating it as a single uninterrupted string of amino acids.
- Protein domains and families provide information about evolutionary and functional relationships between proteins. UniProtKB connects entries with external resources that classify protein domains, families, and conserved regions. Such information can help researchers determine whether a newly studied protein belongs to a known family or contains domains associated with particular biological functions.
- Gene Ontology annotations provide standardized descriptions of molecular function, biological process, and cellular component. By connecting UniProtKB records with Gene Ontology terms, researchers can compare functional information across proteins and organisms using a common vocabulary. These annotations are particularly useful in large-scale functional genomics and enrichment analyses.
- Subcellular location information describes where a protein is found within a cell or biological system. Depending on the available evidence, an entry may provide information about locations such as the nucleus, cytoplasm, mitochondrion, plasma membrane, endoplasmic reticulum, or extracellular space. Localization information can provide important clues about a protein’s biological role.
- Post-translational modification information describes chemical or structural changes that occur to proteins after translation. UniProtKB can contain information about modifications such as phosphorylation, acetylation, glycosylation, lipidation, and other forms of protein processing. Such information can be associated with specific amino acid positions or regions within a protein.
- Protein isoforms can also be represented in UniProtKB when alternative forms of a protein arise through mechanisms such as alternative splicing or other biological processes. Different isoforms may have different sequences, cellular locations, interactions, or biological functions. Therefore, examining isoform information can be important when a gene produces more than one protein form.
- Protein existence evidence provides information about the level of evidence supporting the existence of a protein. Depending on the available information, evidence can include direct protein-level evidence, transcript-level evidence, inference from homology, prediction, or other categories. This distinction is useful because the presence of a predicted protein sequence does not necessarily mean that the corresponding protein has been experimentally observed.
- Cross-references connect UniProtKB entries with other biological resources. A protein entry can provide links to databases and resources covering protein structures, domains, pathways, taxonomy, gene ontology, protein families, genomes, literature, and other biological information. These connections make UniProtKB particularly useful as a starting point for exploring a protein across multiple areas of biological research.
- Scientific references provide a connection between database annotation and published research. UniProtKB entries can contain citations associated with experimental findings, functional information, sequence data, and other aspects of the protein record. Researchers can therefore move from an annotation in UniProtKB to the scientific literature that provides supporting evidence.
- The UniProtKB search system allows users to locate protein records using a wide range of information. A basic free-text search can identify entries containing a particular term, while advanced searches allow users to restrict queries to specific fields and combine conditions using Boolean logic. Users can also filter results according to reviewed or unreviewed status, organism, taxonomy, annotation characteristics, and other properties.
- UniProtKB filters are particularly useful when a search produces a large number of results. Researchers can narrow results according to criteria such as reviewed status, organism, taxonomy, Gene Ontology, protein existence, sequence characteristics, and other available attributes. Filtering allows users to move from a broad protein search toward a more specific set of records suitable for further analysis.
- UniProtKB also supports programmatic access for researchers who need to retrieve protein data automatically. The UniProt REST API allows users to submit queries and retrieve search results or individual entries in machine-readable formats. This makes it possible to incorporate UniProtKB information into bioinformatics pipelines, scripts, databases, and large-scale computational analyses.
- UniProtKB data downloads provide another way to obtain protein information for computational work. Users can download selected datasets from search results, while larger datasets can be obtained through UniProt’s download infrastructure. Reviewed and unreviewed UniProtKB data are available in several formats, including FASTA, XML, and text-based formats.
- Versioning and data history are important when using UniProtKB in reproducible research. UniProtKB is updated regularly, and previous versions of changed entries are preserved in UniSave. This allows researchers to examine earlier versions of records and understand how protein annotations or sequences have changed over time.
- UniProtKB is continuously updated because protein sequence data and biological knowledge are constantly expanding. New sequences are added, existing records are reviewed, annotations are improved, and scientific evidence is incorporated as it becomes available. Consequently, researchers should consider the database release and record version when documenting UniProtKB data used in a study.
- The organization of UniProtKB has also evolved as the amount of available sequence data has grown. UniProt has been reorganizing the unreviewed protein space around Reference Proteomes to improve representation across the diversity of life. Recent changes have affected the composition and size of UniProtKB, while sequences removed from the active UniProtKB collection remain accessible through UniParc.
- UniProtKB is widely used in bioinformatics, molecular biology, genomics, proteomics, structural biology, biotechnology, and biomedical research. A researcher may use it to identify an unknown protein, investigate its function, compare homologous proteins, examine conserved regions, identify domains, investigate protein modifications, obtain sequences for computational analysis, or connect a protein with structural and pathway information.
- For students and beginners, UniProtKB can serve as an important introduction to the way modern biological databases organize scientific knowledge. For experienced researchers, it can serve as a source of high-quality protein data that can be incorporated into computational workflows. Its combination of sequence information, functional annotation, evidence, literature, identifiers, and cross-references makes it one of the most useful resources for studying proteins.
- Understanding how to read a UniProtKB entry is therefore an important skill for anyone working with protein data. A user should learn how to interpret the accession number, protein name, gene name, organism, sequence, function, subcellular location, sequence features, evidence, references, and cross-references. Each part of the entry contributes a different piece of information about the protein.
- In summary, UniProtKB is the central knowledgebase within UniProt for accessing protein sequence and functional information. Its two principal sections, UniProtKB/Swiss-Prot and UniProtKB/TrEMBL, combine expert-reviewed annotation with large-scale computationally processed protein data. Through protein sequences, functional annotation, evidence, sequence features, identifiers, taxonomy, literature, cross-references, search tools, downloads, and programmatic access, UniProtKB provides a comprehensive foundation for protein research. In the next articles, each of these components can be examined individually to understand how UniProtKB records are constructed, searched, interpreted, and used in practical bioinformatics workflows.