Skip to content

Article image
Phylogenomics: Genome-Scale Phylogenetics

May 16, 2026 · Updated: May 25, 2026

Overview

Phylogenomics extends traditional phylogenetics by analyzing data from hundreds or thousands of genes across the genome simultaneously. By leveraging genome-scale information, phylogenomics overcomes the limitations of single-gene phylogenies, which often suffer from insufficient signal, lineage-specific evolutionary rate variation, and stochastic error. The increased data volume dramatically improves statistical power, enabling resolution of even the most challenging and rapid evolutionary radiations. Phylogenomics also exposes incongruences between gene trees and species trees caused by incomplete lineage sorting, gene duplication, and horizontal gene transfer.

Key Concepts

A central challenge is gene tree–species tree reconciliation. Because individual genes can have evolutionary histories that differ from the species tree, phylogenomic methods must account for processes such as incomplete lineage sorting (modeled by the multispecies coalescent) and gene duplication and loss. Concatenation approaches align all genes into a supermatrix, while coalescent-based methods analyze each gene independently and then summarize the results. Orthology assignment, distinguishing orthologs from paralogs, is a critical preprocessing step that relies on accurate genome annotation.

Practical Workflow

A phylogenomics analysis begins with genome sequencing and annotation of the target species. Ortholog identification is the critical first computational step: tools such as OrthoFinder or OrthoMCL cluster proteins into orthogroups based on sequence similarity, using BLAST all-versus-all searches and Markov clustering. For each orthogroup, protein sequences are aligned using MAFFT or Muscle, and the alignments are trimmed with trimAl to remove poorly aligned regions. The researcher then faces a fundamental methodological choice between concatenation and coalescent approaches. In the concatenation approach (the supermatrix method), all gene alignments are joined into a single large alignment, and a phylogenetic tree is inferred using maximum likelihood with RAxML or IQ-TREE under a partitioned model where each gene can have its own substitution model parameters. This approach is computationally efficient but assumes all genes share the same tree topology. The coalescent-based approach instead infers a separate gene tree for each orthogroup using fast tree-building methods, then summarizes the thousands of individual gene trees into a species tree using methods such as ASTRAL, which accounts for incomplete lineage sorting by finding the species tree that maximizes the number of shared quartet trees. Comparing the concatenation and coalescent results is a valuable diagnostic, topological differences often highlight genes under selection or affected by horizontal transfer. Bootstrapping (typically 100 replicates) provides support values at each node, and additional analyses such as tree certainty scores can identify conflict among gene trees.

Applications

Phylogenomics has resolved long-standing debates in deep metazoan phylogeny, plant evolution, and microbial systematics. It is essential for studying adaptive evolution through the identification of positively selected genes across lineages. The field depends on next-generation sequencing to generate the required genome-wide data, building on classical DNA sequencing approaches. Phylogenomic analyses of bacterial genomes have reshaped our understanding of bacterial genetics, revealing extensive horizontal gene transfer and the dynamic nature of prokaryotic genomes.