Key Points
Introduction to metagenomics
- Metagenomics sequences all DNA, giving an unbiased, high-resolution view of community composition and function—beyond what amplicon sequencing can capture.
- Reads profiling delivers fast, reference-based taxonomic and functional counts data for downstream analyses like TaxSEA and integration workflows.
- MAGs can be reconstructed via assembly and binning of reads, revealing uncultured lineages and their metabolic potential.
- Modern sequencing and workflows (short-reads, long-reads, hybrids; nf-core pipelines) improve accuracy, assembly quality, and reproducibility.
Taxon-set enrichment analysis with TaxSEA
- TaxSEA performs taxon-set enrichment analysis on differential abundance results, similar to GSEA.
- Inputs must be species or genus names with corresponding ranks (e.g., log2 fold changes).
- TaxSEA tests whether members of a taxon set are skewed toward one end of the ranked distribution.
- Enrichment results help reveal functional or ecological patterns that may not be visible at the single-taxon level.
Multi-omics integration with DIABLO
- DIABLO integrates multiple omics datasets by identifying correlated and discriminative features across blocks.
- Proper filtering and transformation of each omics dataset are essential for stable multi-omics models.
- Tuning (components and keepX) guides the selection of the most predictive and biologically relevant features.
- DIABLO visualisations (sample plots, arrow plots, circos, CIM) reveal relationships across omics and help interpret multi-omics signatures.