![]()
- TrEMBL is the unreviewed component of the UniProt Knowledgebase (UniProtKB) and provides a very large collection of protein sequence records that are primarily annotated using computational methods. The name TrEMBL originated from the expression “Translated EMBL,” reflecting its historical connection with translated coding sequences from the European Molecular Biology Laboratory (EMBL) nucleotide sequence database. Today, TrEMBL forms the unreviewed section of UniProtKB and works alongside the manually curated UniProtKB/Swiss-Prot section to provide broad coverage of known protein sequences.
- The main purpose of UniProtKB/TrEMBL is to make newly available protein sequence information accessible as quickly and comprehensively as possible. Modern genome and metagenome sequencing projects generate enormous numbers of predicted protein sequences. It would not be practical for every newly generated protein record to undergo manual expert curation before becoming available. TrEMBL therefore uses automated computational annotation to process large numbers of sequences and make them searchable through UniProt.
- The relationship between TrEMBL and UniProtKB is fundamental to understanding the organization of UniProt. UniProtKB has two major sections: UniProtKB/Swiss-Prot, which contains reviewed and manually curated records, and UniProtKB/TrEMBL, which contains unreviewed records. Together, these sections provide both detailed expert-curated protein information and extensive coverage of the rapidly expanding protein sequence universe.
- The term unreviewed protein record refers to a UniProtKB record that has not yet undergone complete manual annotation by UniProt curators. Unreviewed does not necessarily mean that the sequence or computational annotation is incorrect. Instead, it indicates that the record has not received the same level of expert manual review associated with Swiss-Prot. Users should therefore examine the evidence and annotation methods before treating information from an individual TrEMBL record as experimentally established.
- The foundation of TrEMBL is protein sequence data. A large proportion of UniProtKB protein sequences originate from translations of coding sequences submitted to the International Nucleotide Sequence Database Collaboration (INSDC), which includes GenBank, ENA, and DDBJ. These translated sequences can enter the unreviewed protein space and receive computational annotation. This process allows UniProt to keep pace with the rapid expansion of sequence data generated by genome sequencing.
- TrEMBL also accommodates protein sequences from other sources. UniProt integrates sequence information from multiple biological resources and computational analyses so that the Knowledgebase can provide broad coverage across organisms and biological systems. The resulting collection includes proteins from bacteria, archaea, eukaryotes, viruses, and other biological entities.
- Automatic annotation is the defining feature of TrEMBL. Instead of requiring a curator to manually examine every individual protein, computational systems analyze sequences and assign appropriate annotations according to established rules, similarities, domains, conserved regions, and other evidence. Automated annotation makes it possible to process millions of protein sequences at a scale that would be impossible through manual curation alone.
- One important system involved in UniProt annotation is UniRule. UniRule is a collection of annotation rules that are manually curated and then applied automatically to suitable protein sequences. The rules can use information such as sequence signatures, taxonomic conditions, and other characteristics to assign standardized annotations to proteins that meet defined criteria. UniRule therefore combines human expertise in rule creation with automated application at large scale.
- Another system used for automated annotation is ARBA, or Association-Rule-Based Annotator. ARBA uses association rules derived from protein sequence and annotation information to provide computational annotations. Together with other UniProt annotation systems, ARBA contributes to the large-scale annotation of unreviewed protein records.
- Similarity-based annotation is another important concept in TrEMBL. Proteins with similar sequences may share evolutionary relationships and, in some circumstances, similar functions. Computational methods can therefore compare newly generated protein sequences with proteins that already have functional information. However, similarity does not automatically prove identical biological function, so users should consider the evidence and annotation status when interpreting transferred information.
- Protein sequence similarity can help identify homologous proteins, conserved regions, and potential functional relationships. Computational searches and classification systems can detect relationships between a newly predicted protein and proteins that have already been characterized. This is particularly useful when experimental information is available for only a small proportion of proteins in a particular organism or protein family.
- Protein domains and conserved regions provide another source of information for automatic annotation. Computational systems can recognize characteristic sequence patterns or domain signatures and use them to infer possible protein functions or family membership. Connections with resources such as InterPro and related protein classification databases can provide additional context for TrEMBL records.
- Gene Ontology annotations can also be associated with proteins through computational and evidence-based methods. Gene Ontology provides standardized terms describing molecular function, biological process, and cellular component. These annotations allow proteins in TrEMBL to be incorporated into large-scale functional analyses, although users should consider the evidence code and method behind each annotation.
- Sequence features can be predicted or inferred for TrEMBL proteins. These may include signal peptides, transmembrane regions, domains, active-site candidates, binding regions, repeats, and other characteristics. Computational predictions can be extremely useful for investigating newly identified proteins, particularly when experimental characterization is not yet available.
- Protein function prediction is one of the major applications of TrEMBL annotation. When a protein has not been experimentally characterized, computational methods can provide hypotheses about its possible function. These predictions can guide further research and help researchers prioritize proteins for experimental investigation. However, predicted function should be distinguished from experimentally demonstrated function.
- The distinction between prediction and experimental evidence is therefore essential when using TrEMBL. A computational annotation can provide a valuable hypothesis, but it should not automatically be interpreted as proof. Researchers should examine the evidence associated with an annotation and, when necessary, consult the scientific literature or reviewed Swiss-Prot records for additional support.
- Evidence attribution allows users to investigate how annotations were assigned. UniProt distinguishes different evidence types and provides information that helps users understand whether an annotation is supported experimentally, inferred computationally, transferred from another protein, or generated through another annotation process. Understanding evidence is particularly important when using unreviewed records in research.
- The enormous size of TrEMBL is directly related to the growth of modern sequencing technologies. Genome sequencing, metagenomics, transcriptomics, and other high-throughput approaches can generate huge numbers of predicted proteins. TrEMBL provides a way to capture this information within the UniProt ecosystem without requiring each protein to wait for manual curation before becoming available.
- Metagenomic protein sequences are particularly relevant to the expansion of the unreviewed protein space. Metagenomic sequencing can reveal genetic material from microbial communities containing organisms that may be difficult or impossible to culture under laboratory conditions. Computational translation of these sequences can generate large numbers of predicted proteins, many of which may have limited experimental characterization.
- The broad taxonomic coverage of TrEMBL makes it useful for studying protein diversity across organisms. Researchers can search for proteins from particular species, genera, families, or broader taxonomic groups. This can support comparative genomics, evolutionary studies, environmental microbiology, microbial genomics, and investigations of protein families.
- Taxonomic information is therefore an important component of TrEMBL records. Each protein sequence is associated with an organism or taxonomic context, allowing researchers to investigate how proteins are distributed across the tree of life. Taxonomic filtering can also help reduce large search results to a biologically relevant group.
- A major advantage of TrEMBL is sequence coverage. Whereas manual curation can provide very detailed information for selected proteins, automatic annotation allows UniProt to represent a much larger proportion of the known protein sequence space. This broad coverage is particularly valuable for organisms and protein families for which relatively little experimental characterization is available.
- The relationship between TrEMBL and Swiss-Prot can therefore be understood as complementary rather than competitive. Swiss-Prot emphasizes depth, expert review, and manual curation, while TrEMBL emphasizes scale and comprehensive coverage through computational processing. Together, they allow UniProtKB to provide both high-quality reviewed information and a much broader collection of newly available protein sequences.
- Some TrEMBL records can eventually become reviewed Swiss-Prot records. When sufficient biological information becomes available and a protein is selected for manual curation, UniProt curators can review the available evidence and update the record. The transition from an unreviewed to a reviewed status represents an important improvement in the level of manual annotation associated with the protein.
- The conversion of a TrEMBL record into Swiss-Prot does not necessarily mean that the underlying protein sequence has changed. Instead, the record can receive substantial additional biological interpretation through manual curation. Protein names, functional descriptions, sequence features, references, evidence, and other annotations may be refined as curators incorporate scientific knowledge.
- UniProtKB accession numbers provide identifiers for TrEMBL protein records just as they do for reviewed records. These identifiers allow researchers to retrieve and reference individual protein entries. When using TrEMBL data in computational studies, accession numbers can be particularly useful for maintaining links between sequence datasets and UniProt records.
- UniProtKB entry names provide another way of identifying protein records. Entry names can be useful for human-readable identification, while accession numbers are generally more suitable as stable database identifiers. Researchers should distinguish between descriptive protein names, gene names, entry names, and accession numbers when working with UniProt data.
- UniProt search allows researchers to locate TrEMBL records using protein names, gene names, accession numbers, organisms, taxonomy, functional terms, sequence characteristics, and other criteria. Search filters can then be used to restrict results to unreviewed entries or particular biological categories. This makes it possible to work with TrEMBL datasets without manually examining individual records.
- UniProt filters are especially important because of the enormous number of unreviewed records. Researchers can narrow results using criteria such as organism, taxonomy, reviewed status, protein existence, Gene Ontology, sequence properties, and other attributes. Effective filtering can make large TrEMBL datasets much easier to analyze.
- TrEMBL data can also be accessed through the UniProt REST API. Programmatic access allows researchers to retrieve large collections of protein records and incorporate them into computational workflows. This is particularly useful in genome annotation, comparative genomics, sequence analysis, proteomics, and other bioinformatics applications.
- UniProt data downloads provide another method for working with TrEMBL information. Researchers can retrieve sequences and annotations in machine-readable formats and use them with sequence-analysis software, databases, scripts, and bioinformatics pipelines. Large-scale data access is particularly important when studying entire proteomes or large protein families.
- TrEMBL is also connected to other UniProt resources, including UniRef and UniParc. UniRef groups similar sequences into clusters to reduce redundancy and facilitate efficient sequence analysis, while UniParc provides an archive of protein sequences and their historical database records. These resources complement the active protein knowledge maintained in UniProtKB.
- Cross-references connect TrEMBL records with other biological databases and resources. Depending on the protein and available information, links can lead to databases covering protein domains, structures, pathways, taxonomy, Gene Ontology, genomes, protein families, and scientific literature. These connections allow researchers to combine TrEMBL information with specialized resources.
- Protein structure prediction can be particularly useful for unreviewed proteins. When experimental structural information is unavailable, computational structure resources can provide predicted models that help researchers investigate possible folds and functional regions. However, predicted structure should be distinguished from experimentally determined structure, just as predicted function should be distinguished from experimentally demonstrated function.
- TrEMBL can be particularly valuable in genome annotation. When a newly sequenced genome contains thousands of predicted coding regions, researchers need a way to identify and characterize the corresponding proteins. Computational comparison with existing UniProtKB records and other resources can help assign preliminary functional descriptions and identify proteins that may require further investigation.
- TrEMBL also plays an important role in comparative genomics. Researchers can compare proteins from different organisms and investigate conserved sequences, protein families, domain organization, and evolutionary relationships. The extensive taxonomic coverage of TrEMBL makes it particularly useful when studies involve organisms for which relatively few proteins have been manually reviewed.
- In proteomics, TrEMBL sequences can be useful for identifying proteins from experimental mass-spectrometry data, especially when studying organisms or samples with incomplete reference proteomes. However, researchers should carefully evaluate the quality and redundancy of the underlying sequence database because unreviewed and computationally predicted proteins can influence protein identification and downstream interpretation.
- The use of TrEMBL data in machine learning and computational biology is another important application. Large collections of protein sequences and annotations can be used to train or evaluate computational models for protein classification, function prediction, sequence representation, and other tasks. Researchers must nevertheless consider annotation quality, redundancy, data leakage, and evidence levels when constructing machine-learning datasets.
- Because TrEMBL is automatically annotated, data quality assessment is important. Computational annotations can contain errors or propagate incorrect assumptions when predictions are transferred between related proteins. Users should therefore avoid treating every annotation as equally reliable and should consider reviewed status, evidence codes, sequence similarity, experimental literature, and other supporting information.
- The unreviewed status of a TrEMBL record should consequently be interpreted carefully. It does not indicate that the protein is unimportant or that the information is necessarily wrong. It primarily indicates that the record has not undergone complete manual review. Many unreviewed proteins are scientifically valuable and may contain useful computational annotations that can guide future experiments.
- TrEMBL is continuously affected by database updates and releases. As new sequence data become available, additional records can be incorporated, while existing records can receive updated annotations or be reorganized as UniProt improves its data-processing and reference-proteome strategies. Researchers performing reproducible analyses should therefore record the UniProt release or dataset version used.
- The organization of the unreviewed protein space has evolved as UniProt has sought to manage the rapid growth of sequence databases. UniProt has increasingly emphasized Reference Proteomes and improved ways of representing the enormous diversity of unreviewed proteins. These developments are intended to make the database more useful while controlling redundancy and improving the representation of biological diversity.
- TrEMBL is especially important for organisms and proteins that have not yet received extensive experimental study. A newly sequenced bacterial genome, for example, may contain thousands of predicted proteins for which no experimental functional information exists. TrEMBL can provide immediate computationally derived information about these proteins and make them available for further analysis.
- For researchers, one of the most useful approaches is to treat TrEMBL annotation as a starting point for investigation. A predicted function can suggest what experiments should be performed, a predicted domain can suggest a molecular role, and sequence similarity can identify potentially related proteins. Researchers can then investigate the available evidence and determine whether the proposed biological interpretation is sufficiently supported.
- For students learning bioinformatics, TrEMBL provides an important example of how modern biological databases manage the enormous volume of information produced by high-throughput sequencing. It demonstrates why computational annotation is necessary and why users must understand the difference between prediction, inference, and experimental validation.
- In summary, UniProtKB/TrEMBL is the unreviewed, computationally annotated component of the UniProt Knowledgebase. It provides extensive coverage of protein sequences generated through modern sequencing projects and uses automated approaches such as UniRule, ARBA, sequence similarity, domain analysis, and other computational methods to provide preliminary biological information. Its enormous scale complements the manually curated Swiss-Prot collection and allows UniProt to represent a much broader portion of known protein sequence diversity.
- Understanding TrEMBL is therefore essential for understanding how UniProt manages the rapidly expanding protein sequence universe. While Swiss-Prot provides manually reviewed and carefully curated protein knowledge, TrEMBL provides the large-scale computational layer that captures newly available sequences and makes them accessible for research. In the following articles, the individual aspects of TrEMBL—including automatic annotation, UniRule, ARBA, evidence codes, computational function prediction, reviewed versus unreviewed records, Reference Proteomes, and the transition from TrEMBL to Swiss-Prot—can be explored in greater detail.