Key Points

Background


  • Captivity can alter diet, exposure and behaviour — all of which may reshape the gut microbiome.
  • The dataset contains a small, balanced subset (5 captive, 5 wild) suitable for teaching and demonstrating methods.
  • Use byobu-screen to keep long‑running commands alive across disconnections.
  • Symlinks point to /mnt/shared_data to conserve storage and keep everyone working from the same files.

Importing, cleaning and quality control of the data


  • Confirm input format (CasavaOneEightSingleLanePerSampleDirFmt) when importing.
  • Primer removal is essential; untrimmed primers can disrupt denoising and downstream inference.
  • Inspect interactive quality plots and consider read overlap (~50 bp recommended) when setting truncation lengths for paired‑end merging.
  • Use dada2 denoise‑paired with appropriate pooling (here ‘pseudo’) and monitor read retention in denoising_stats outputs.

Taxonomic Analysis


  • Classifier must be trained for the same primer/region used in the dataset; mismatched classifiers give poor results.
  • The classify‑sklearn step can be memory intensive — precomputed classifications are provided for the workshop.
  • Retain unassigned ASVs (they may still be biologically meaningful), but exclude obvious host/plant contaminants.
  • Low‑abundance filtering is generally unnecessary with the ASV paradigm unless specific reasons exist.

Build a phylogenetic tree


  • Phylogenetic distances are required for phylogenetic alpha and beta metrics (Faith’s PD, UniFrac).
  • Use align‑to‑tree‑mafft‑fasttree for a reproducible single‑threaded workflow.
  • Keep track of the unrooted and rooted tree artefacts for downstream use.

Basic visualisations and statistics


  • Explore taxonomic composition at multiple taxonomic levels and by metadata categories (e.g. Captivity).
  • Choose sampling/depth parameters informed by the feature‑table summary and rarefaction curves to balance sample retention and depth.
  • Unweighted and weighted UniFrac capture different aspects (presence/absence vs abundance) of community dissimilarity.
  • ANCOM‑BC2 is compositionally aware and suitable for differential abundance while applying bias correction.

Exporting data for further analysis in R


  • Exported artefacts include: tree (Newick), feature‑table.biom, taxonomy.tsv and representative sequences if required.
  • Use biom convert to produce TSV tables and sed (or equivalent) to remove extra header lines.
  • Confirm consistent ASV order between taxonomy and table files before importing into phyloseq or other R packages.

Extra Information


  • Training a classifier requires region‑matched reference reads and taxonomy; forward/reverse primers must be provided when extracting reads.
  • Classifier training and full classify‑sklearn runs can be computationally demanding — Consider precomputed classifiers or higher memory/CPU resources for large datasets.