![]()
- Understanding what a protein does is one of the central goals of protein annotation, and UniProt provides a structured way to describe protein function using information from experimental studies, scientific literature, sequence analysis, computational inference, and expert curation. Rather than treating protein function as a single label, UniProt can describe several aspects of biological activity, including molecular function, catalytic activity, biological role, interactions, localization, pathways, and functional domains.
- Within UniProtKB, the function of a protein is primarily communicated through the annotation associated with an individual protein entry. A protein entry can contain a functional description explaining the known or proposed role of the protein, together with additional annotations that provide more specific information about its molecular activity and biological context. These annotations allow users to move from a raw amino acid sequence toward a biological interpretation.
- The UniProt function annotation is designed to summarize what is known about the protein’s biological role. For example, an entry may describe a protein as an enzyme involved in a particular metabolic reaction, a receptor involved in signal transduction, a transporter responsible for moving molecules across a membrane, or a regulatory protein controlling another biological process. The level of detail depends on the available evidence.
- Protein function can be considered at several biological levels. At the molecular level, a protein may bind a particular molecule, catalyze a reaction, recognize a substrate, or interact with another protein. At the cellular or organismal level, the same protein may participate in a signaling pathway, metabolic process, developmental process, immune response, or other biological function. UniProt can connect these different levels through its annotation system.
- For enzymes, UniProt can provide detailed catalytic activity information. This describes the biochemical reaction performed by an enzyme and may be accompanied by an EC number when the enzyme has been assigned an appropriate Enzyme Commission classification. Catalytic annotations can also be connected to substrates, products, cofactors, catalytic residues, and other biochemical characteristics.
- Enzyme function annotation can therefore be much more informative than simply stating that a protein is an enzyme. An entry may identify the reaction catalyzed, describe substrate specificity, indicate whether particular cofactors are required, and identify residues that participate directly in catalysis. This information is particularly useful when studying metabolic pathways or characterizing newly identified enzymes.
- UniProt also describes molecular function using controlled terminology and structured annotations. Molecular function refers to what a protein or protein-associated molecular entity does at the molecular level, such as binding, catalysis, transport, or regulation. This information can be integrated with Gene Ontology (GO) annotations, which provide standardized terms for describing molecular functions.
- Gene Ontology annotation provides a standardized framework that separates molecular function from biological process and cellular component. A protein can therefore be associated with a molecular activity while also being linked to the biological processes in which that activity participates and the cellular location where it occurs. This makes functional information easier to compare across proteins and organisms.
- A protein’s biological process describes the larger biological activity in which it participates. For example, a protein may contribute to DNA replication, protein synthesis, carbohydrate metabolism, cell signaling, immune response, or cellular transport. Biological-process annotations provide context that cannot always be inferred from a molecular function alone.
- Subcellular location also contributes to understanding protein function. A protein found in the nucleus may have a very different biological role from a protein located in the mitochondrial inner membrane or extracellular space. UniProt can describe cellular and subcellular localization and, when appropriate, connect this information with sequence features and experimental evidence.
- Protein localization can sometimes be inferred from sequence characteristics. Signal peptides, transmembrane regions, organelle-targeting sequences, and other features can provide clues about where a protein is transported or embedded. UniProt can annotate these sequence features separately while using them as part of the broader functional interpretation of the protein.
- Protein domains are another important component of functional description. Many domains represent conserved structural or functional units that occur in multiple proteins. Identifying a particular domain can provide evidence that a protein belongs to a functional family or possesses a particular molecular capability.
- Protein families provide evolutionary and functional context. Proteins that share significant sequence similarity and evolutionary relationships often retain related functions, although sequence similarity alone does not always establish identical biological activity. UniProt uses family and domain information as part of its broader annotation framework rather than treating similarity as automatic proof of function.
- Active-site annotation provides more precise information for proteins whose function depends on particular amino acid residues. UniProt can identify residues that participate directly in catalysis when suitable evidence exists. This helps researchers connect the overall functional description with specific positions in the protein sequence.
- Binding-site annotation similarly identifies residues or regions involved in binding molecules such as substrates, nucleic acids, cofactors, metal ions, or other proteins. A binding site can provide important evidence about how a protein performs its molecular function and can also be useful when studying mutations that affect activity.
- Cofactor annotation describes molecules or ions required for protein activity. Some enzymes depend on metal ions, nucleotide-derived cofactors, vitamins, or other chemical groups. Recording these requirements helps researchers understand the biochemical mechanism represented by the functional annotation.
- UniProt can also describe protein-protein interactions when suitable information is available. A protein may perform its biological role by interacting with one or more partner proteins, and these interactions can be essential for signaling, complex formation, regulation, transport, or enzymatic activity. Interaction information therefore adds another layer to a protein’s functional description.
- Pathway annotation places individual proteins within larger biochemical or cellular networks. A protein may participate in a metabolic pathway, signaling cascade, DNA repair pathway, immune pathway, or other biological system. UniProt can connect protein entries to pathway resources through cross-references, allowing users to explore function at the systems level.
- The literature evidence behind an annotation is particularly important. Functional descriptions may be supported by experimental studies reported in scientific publications. UniProt links appropriate annotations to literature references so that researchers can investigate the original evidence and understand how a particular functional conclusion was established.
- UniProt distinguishes between different types of protein function evidence. Some annotations are supported directly by experimental observations, while others are inferred computationally or transferred from related proteins. This distinction is essential because a predicted function and an experimentally demonstrated function do not necessarily have the same evidential strength.
- In UniProtKB/Swiss-Prot, expert curators examine scientific literature and other sources when creating and updating reviewed protein entries. Manual curation allows functional information to be evaluated in its biological context and represented consistently within the database. Curators can incorporate experimental findings, resolve terminology, select appropriate descriptions, and connect functional statements to supporting evidence.
- In UniProtKB/TrEMBL, functional information is primarily generated through computational approaches because of the enormous number of unreviewed protein sequences. Automatic protein annotation can use sequence similarity, protein families, rules, conserved regions, and other computational evidence to assign functional information. Such annotations can be extremely useful but should be interpreted with awareness of their evidence and reviewed status.
- UniRule is one of the systems used by UniProt for large-scale computational annotation. It uses predefined rules based on protein characteristics and existing biological knowledge to assign appropriate annotations to sequences satisfying particular conditions. This helps maintain consistency when similar proteins occur across many organisms.
- ARBA provides another computational annotation approach used by UniProt. It derives annotation rules from associations observed within protein data and can apply those rules to suitable sequences. Together, UniRule and ARBA help extend functional information across the large unreviewed protein space represented in UniProtKB.
- Sequence similarity-based annotation is particularly important for newly identified proteins. If an unknown sequence is closely related to a protein with a well-established function, that relationship may provide evidence for a similar function. However, researchers should distinguish between functional similarity and proven functional identity, especially when proteins have diverged or acquired different substrate specificities.
- UniProt may also use conserved residues and conserved sequence patterns when describing function. If particular residues are strongly conserved across a protein family and are known to be important for activity, their presence in another sequence can provide supporting evidence for a related function. These observations can complement domain, family, and similarity-based analyses.
- Post-translational modifications (PTMs) can influence protein function and may therefore appear alongside functional information. Phosphorylation, glycosylation, acetylation, ubiquitination, lipidation, and other modifications can alter protein activity, localization, stability, or interactions. UniProt can record known modification sites and connect them to the corresponding protein sequence.
- Protein processing can also affect function. Some proteins are synthesized as precursors and require cleavage before becoming biologically active. UniProt can annotate signal peptides, propeptides, mature chains, and other processed regions when appropriate evidence exists. This helps explain why the mature functional protein may differ from the initially translated sequence.
- Protein isoforms can have different functional properties despite originating from the same gene. Alternative splicing or other mechanisms can produce protein variants with different domains, localization patterns, interaction partners, or biological activities. UniProt can represent isoform-specific information when sufficient evidence is available.
- Disease-related function is another important aspect of functional annotation for human proteins. A protein may be associated with a disease because its normal function is disrupted, its expression or localization changes, or particular sequence variants alter its activity. UniProt can connect functional information with disease and variant annotations when supported by appropriate evidence.
- Sequence variants can provide clues about the molecular basis of altered protein function. A substitution in an active site, binding region, transmembrane segment, or other functionally important feature may affect protein activity or stability. UniProt can associate relevant variants with functional or disease information where appropriate.
- An important distinction is that protein function prediction is not equivalent to experimentally establishing protein function. Computational prediction can provide valuable hypotheses, particularly for proteins that have not been characterized experimentally, but predictions should be evaluated according to their evidence and annotation status. This distinction is especially relevant when working with unreviewed UniProtKB entries.
- The reviewed status of an entry provides useful context when interpreting functional descriptions. Reviewed UniProtKB entries in Swiss-Prot have undergone expert manual curation, while unreviewed entries in TrEMBL have primarily computational annotations. Researchers should therefore consider both the functional statement and the evidence supporting it rather than treating every annotation as equally established.
- Functional annotations can also evolve as scientific knowledge improves. A protein may initially be described using a broad functional term and later receive a more precise description after new experiments reveal its substrate, molecular mechanism, localization, or biological role. UniProt updates its entries as new evidence becomes available.
- For large-scale analyses, UniProt functional information can be retrieved using the UniProt REST API, downloadable datasets, and web-based search tools. Researchers can use these resources to obtain functional descriptions, GO terms, enzyme classifications, sequence features, evidence information, and other annotations for individual proteins or large collections of sequences.
- A useful way to interpret a UniProt functional description is to read it together with the evidence, sequence features, domains, localization, literature references, and cross-references. The function field provides the central biological summary, while the surrounding annotation explains why that interpretation is reasonable and how it relates to the underlying protein sequence.
- When analyzing an unfamiliar protein, researchers should therefore avoid relying on the protein name alone. A stronger interpretation comes from combining the UniProt function annotation, evidence attribution, sequence similarity, domains, catalytic or binding sites, Gene Ontology terms, localization, literature, and other relevant annotations. This layered approach provides a more reliable understanding of what the protein is likely to do.
- Ultimately, UniProt describes protein function as a combination of biological statements supported by different types of evidence. Rather than reducing a protein to a single functional label, UniProt connects molecular activity with biological processes, cellular location, sequence features, domains, interactions, pathways, modifications, variants, and literature. This integrated approach makes UniProtKB one of the most useful resources for interpreting protein sequences and understanding their biological roles.