Key Points

Introduction to metagenomics


  • Metagenomics sequences all DNA, giving an unbiased, high-resolution view of community composition and function—beyond what amplicon sequencing can capture.
  • Reads profiling delivers fast, reference-based taxonomic and functional counts data for downstream analyses like TaxSEA and integration workflows.
  • MAGs can be reconstructed via assembly and binning of reads, revealing uncultured lineages and their metabolic potential.
  • Modern sequencing and workflows (short-reads, long-reads, hybrids; nf-core pipelines) improve accuracy, assembly quality, and reproducibility.

Taxon-set enrichment analysis with TaxSEA


  • TaxSEA performs taxon-set enrichment analysis on differential abundance results, similar to GSEA.
  • Inputs must be species or genus names with corresponding ranks (e.g., log2 fold changes).
  • TaxSEA tests whether members of a taxon set are skewed toward one end of the ranked distribution.
  • Enrichment results help reveal functional or ecological patterns that may not be visible at the single-taxon level.

Multi-omics integration with DIABLO


  • DIABLO integrates multiple omics datasets by identifying correlated and discriminative features across blocks.
  • Proper filtering and transformation of each omics dataset are essential for stable multi-omics models.
  • Tuning (components and keepX) guides the selection of the most predictive and biologically relevant features.
  • DIABLO visualisations (sample plots, arrow plots, circos, CIM) reveal relationships across omics and help interpret multi-omics signatures.