Skip to content

Article image
Protein Identification by Mass Spectrometry

May 16, 2026 · Updated: May 25, 2026

Overview

Protein identification by mass spectrometry is the cornerstone of proteomics. The fundamental workflow involves digesting proteins into peptides using a protease such as trypsin, measuring the masses of the intact peptides (MS1), and then fragmenting selected peptides to generate tandem mass spectra (MS/MS) that reveal their amino acid sequence. The resulting spectra are matched against theoretical spectra derived from protein sequence databases to assign identities. This bottom-up or shotgun approach is highly scalable and can identify thousands of proteins from complex mixtures in a single experiment.

Methods

Peptide mass fingerprinting (PMF) identifies proteins by matching the list of experimentally measured peptide masses against the theoretical peptide masses calculated from a database. PMF works best for simple protein mixtures or purified proteins separated by techniques such as SDS-PAGE. Tandem MS (MS/MS) identification fragments individual peptides to produce sequence-informative spectra. Search engines compare the observed fragment ion series, primarily b- and y-ions, against predicted series from candidate peptides. De novo sequencing infers the peptide sequence directly from the spectrum when no database match is found, using the mass differences between consecutive fragment ions to derive the sequence.

Practical Protocol

A typical protein identification workflow starts with a protein mixture, such as an immunoprecipitated complex, an SDS-PAGE gel band, or a capillary gel electrophoresis fraction. Proteins are reduced with 10 mM dithiothreitol (DTT) at 56°C for 30 minutes, alkylated with 55 mM iodoacetamide in the dark for 30 minutes, and digested overnight with trypsin at a 1:50 enzyme-to-substrate ratio at 37°C. The resulting peptides are desalted using C18 ZipTips and loaded onto a nanoLC system coupled to an Orbitrap mass spectrometer. The instrument operates in data-dependent acquisition mode: a full MS1 scan from 350–1,600 m/z at 70,000 resolution identifies the top 15 most intense precursor ions, which are sequentially isolated and fragmented by higher-energy collisional dissociation (HCD) at 27% normalized collision energy to produce MS/MS spectra. Fragment spectra are searched against a protein database using Mascot or SEQUEST. Each peptide-spectrum match is scored based on the correlation between experimental fragment ions and theoretical b- and y-ion series. A target-decoy search estimates the FDR, and matches are filtered at 1% FDR at both peptide and protein levels. Protein grouping collapses peptides that map to multiple proteins according to parsimony. An example real-world application: identification of the human nuclear pore complex composition via biochemical fractionation and mass spectrometry revealed over 30 nucleoporins, many previously uncharacterized, providing a comprehensive molecular picture of nuclear transport. Another application is bacterial identification by MALDI-TOF MS in clinical microbiology, where ribosomal protein profiles serve as taxonomic fingerprints for species-level identification within minutes rather than days.

Applications

Protein identification is applied throughout molecular biology and clinical research. It confirms the identity of purified proteins obtained through protein extraction and purification, characterizes components of proteomics and mass spectrometry experiments, and identifies protein complexes co-purified with a bait protein. In microbiology, it is used to identify bacterial species through mass spectrometry-based proteotyping, complementing genomic approaches.