Key Points
Background
- Captivity can alter diet, exposure and behaviour — all of which may reshape the gut microbiome.
- The dataset contains a small, balanced subset (5 captive, 5 wild) suitable for teaching and demonstrating methods.
- Use byobu-screen to keep long‑running commands alive across disconnections.
- Symlinks point to /mnt/shared_data to conserve storage and keep everyone working from the same files.
Importing, cleaning and quality control of the data
- Confirm input format (CasavaOneEightSingleLanePerSampleDirFmt) when importing.
- Primer removal is essential; untrimmed primers can disrupt denoising and downstream inference.
- Inspect interactive quality plots and consider read overlap (~50 bp recommended) when setting truncation lengths for paired‑end merging.
- Use dada2 denoise‑paired with appropriate pooling (here ‘pseudo’) and monitor read retention in denoising_stats outputs.
Taxonomic Analysis
- Classifier must be trained for the same primer/region used in the dataset; mismatched classifiers give poor results.
- The
classify‑sklearnstep can be memory intensive — precomputed classifications are provided for the workshop. - Retain unassigned ASVs (they may still be biologically meaningful), but exclude obvious host/plant contaminants.
- Low‑abundance filtering is generally unnecessary with the ASV paradigm unless specific reasons exist.
Build a phylogenetic tree
- Phylogenetic distances are required for phylogenetic alpha and beta metrics (Faith’s PD, UniFrac).
- Use align‑to‑tree‑mafft‑fasttree for a reproducible single‑threaded workflow.
- Keep track of the unrooted and rooted tree artefacts for downstream use.
Basic visualisations and statistics
- Explore taxonomic composition at multiple taxonomic levels and by metadata categories (e.g. Captivity).
- Choose sampling/depth parameters informed by the feature‑table summary and rarefaction curves to balance sample retention and depth.
- Unweighted and weighted UniFrac capture different aspects (presence/absence vs abundance) of community dissimilarity.
- ANCOM‑BC2 is compositionally aware and suitable for differential abundance while applying bias correction.
Exporting data for further analysis in R
- Exported artefacts include: tree (Newick), feature‑table.biom, taxonomy.tsv and representative sequences if required.
- Use biom convert to produce TSV tables and sed (or equivalent) to remove extra header lines.
- Confirm consistent ASV order between taxonomy and table files before importing into phyloseq or other R packages.
Extra Information
- Training a classifier requires region‑matched reference reads and taxonomy; forward/reverse primers must be provided when extracting reads.
- Classifier training and full classify‑sklearn runs can be computationally demanding — Consider precomputed classifiers or higher memory/CPU resources for large datasets.