Overview
UniProt (Universal Protein Resource) is the world’s leading repository of protein sequence and functional information. It is maintained by a consortium that includes the European Bioinformatics Institute (EMBL-EBI), the Swiss Institute of Bioinformatics (SIB), and the Protein Information Resource (PIR). UniProt integrates data from genome sequencing projects, literature curation, and computational predictions to provide a single, comprehensive view of each known protein. Its three main components, UniProtKB, UniRef, and UniParc, serve different analytical needs.
Key Concepts
UniProt Knowledgebase (UniProtKB) is divided into two sections: Swiss-Prot contains manually curated, reviewed entries with detailed annotation derived from literature; TrEMBL contains computationally analyzed entries awaiting manual review. UniRef clusters protein sequences at various identity levels (100%, 90%, 50%) to reduce redundancy and accelerate sequence similarity searches. UniParc is a comprehensive, non-redundant archive that tracks all protein sequences from the major source databases. Each UniProtKB entry includes the sequence, function information, subcellular location, post-translational modifications, domains and sites, and cross-references to other databases.
Applications
UniProt serves as the primary protein reference for most bioinformatics workflows. Researchers use it to retrieve sequences for protein structure prediction, to identify proteins in proteomics and mass spectrometry experiments via database searching, and to look up functional annotations such as enzyme classification and nomenclature for metabolic pathway studies.
Practical Protocol
To retrieve a specific protein sequence from UniProtKB, navigate to uniprot.org and enter the gene name, protein name, or accession (e.g., TP53 or P04637) in the search bar. The results page displays matching entries with annotation scores. Click an entry to view the full record: the sequence appears in the “Sequence” section with FASTA format available via a single click. For batch retrieval, use the “Retrieve/ID mapping” tool, paste a list of accessions or gene names, select the source database (e.g., UniProtKB AC/ID), and choose the target format (FASTA, tab-separated, or GFF). The tool also maps identifiers across 200+ external databases. For conserved domain analysis, take the retrieved FASTA sequence and submit it to InterProScan (integrated within UniProt) or NCBI CD-Search. The results will highlight conserved domains, active sites, and binding regions, providing functional insight even for uncharacterized proteins. For example, submitting a hypothetical protein from a metagenomic sample to UniProt BLAST may reveal homology to a known lipase, with InterProScan confirming the alpha/beta hydrolase fold and catalytic triad, immediately suggesting its biochemical function. This combined retrieval-and-analysis workflow transforms UniProt from a static sequence repository into a discovery platform for functional annotation of novel proteins. Real-world use: during the COVID-19 pandemic, researchers used UniProt to rapidly access the SARS-CoV-2 spike protein sequence and map its receptor-binding domain, accelerating structural studies and vaccine design, demonstrating how UniProt serves as the critical starting point for rapid pandemic response research.