All in One View

Content from Background


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • Why is the fat‑tailed dunnart a useful model for microbiome studies?
  • What are the expected experimental differences between captive and wild animals?
  • Why use a byobu / screen session on a remote instance?
  • What are symbolic links and why are they used here?

Objectives

  • Place the dataset in ecological and conservation context.
  • Relate host ecology and sample provenance to interpretation of microbiome results.
  • Launch and reconnect to a persistent byobu‑screen session.
  • Create symbolic links to shared tutorial data to avoid redundant copies.

What is the influence of captivity on gut microbiota of the fat-tailed dunnart?

The Players


dunnart (Photo credit: Emily Scicluna)

  • Fat-tailed dunnart Sminthopsis crassicaudata - a species of mouse-like marsupial in the family Dasyuridae, which includes quolls, the Tasmanian devil, and the extinct Thylacine. There are 10 samples in this dataset (This data is a subset from a larger experiment); 5 faecal samples each from captive and wild fat-tailed dunnarts.

The Study


Indigenous microbial communities (microbiota) play critical roles in host health. Small marsupials, such as the fat-tailed dunnart, are increasingly used as model systems to understand how environmental conditions shape host-associated microbiomes. Transitions between wild and captive environments can substantially alter diet, behaviour, and microbial exposure, providing a natural framework to investigate microbiome restructuring and its potential consequences for host physiology and health. Here, we characterise the gut microbiome of wild and captive fat-tailed dunnarts to assess how captivity influences microbial community composition. This dataset represents a subset of a larger experimental framework examining microbiome-mediated effects on host function and conservation outcomes.

QIIME 2 Analysis platform


Caution

The version used in this workshop is qiime2-2026.1. Other versions of QIIME2 may result in minor differences in results.

Quantitative Insights Into Microbial Ecology 2 (QIIME 2™) is a next-generation microbiome bioinformatics platform that is extensible, free, open source, and community developed. It allows researchers to:

  • Automatically track analyses with decentralised data provenance
  • Interactively explore data with beautiful visualisations
  • Easily share results without QIIME 2 installed
  • Plugin-based system — researchers can add in tools as they wish

Viewing QIIME2 visualisations

Callout

In order to use QIIME2 View to visualise your files, you will need to use a Google Chrome or Mozilla Firefox web browser (not in private browsing). For more information, click here.

As this workshop is being run on a remote Nectar Instance, you will need to download the visual files (*.qzv) to your local computer and view them in QIIME2 View (q2view).

Callout

We will be doing this step multiple times throughout this workshop to view visualisation files as they are generated.


Alternatively, if you have QIIME2 installed and are running it on your own computer, you can use qiime tools view to view the results from the command line (e.g. qiime tools view filename.qzv). qiime tools view opens a browser window with your visualization loaded in it. When you are done, you can close the browser window and press ctrl-c on the keyboard to terminate the command.

Key Points
  • Captivity can alter diet, exposure and behaviour — all of which may reshape the gut microbiome.
  • The dataset contains a small, balanced subset (5 captive, 5 wild) suitable for teaching and demonstrating methods.
  • Use byobu-screen to keep long‑running commands alive across disconnections.
  • Symlinks point to /mnt/shared_data to conserve storage and keep everyone working from the same files.

Content from Importing, cleaning and quality control of the data


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • How do you import demultiplexed paired‑end FASTQ files into a QIIME 2 artefact?
  • Why must primers be removed prior to denoising?
  • What diagnostics should you inspect before selecting DADA2 truncation parameters?

Objectives

  • Import Casava‑formatted demultiplexed paired‑end reads into a single SampleData artefact.
  • Remove primers with cutadapt while handling degenerate bases and quality trimming.
  • Generate demux quality summaries and choose appropriate DADA2 truncation lengths to balance read quality and merge overlap.

Import data


These samples were sequenced on a single Illumina NextSeq run at the Walter and Eliza Hall Institute (WEHI), Melbourne, Australia. Data from WEHI came as paired-end, demultiplexed, unzipped *.fastq files with adapters still attached. Following the QIIME2 importing tutorial, this is the Casava One Eight format. The files have been renamed to satisfy the Casava format as SampleID_FWDXX-REVXX_L001_R[1 or 2]_001.fastq e.g. CTRLA_Fwd04-Rev25_L001_R1_001.fastq.gz. The files were then zipped (.gzip).

Here, the data files (two per sample i.e. forward and reverse reads R1 and R2 respectively) will be imported and exported as a single QIIME 2 artefact file. These samples are already demultiplexed (i.e. sequences from each sample have been written to separate files), so a metadata file is not initially required.

Callout

To check the input syntax for any QIIME2 command, enter the command, followed by --help e.g. qiime tools import --help

Caution

If you haven’t already done so, make sure you are running the workshop in byobu-screen and have created the symbolic links to the workshop data.

Start by making a new directory analysis to store all the output files from this tutorial. In addition, we will create a subdirectory called seqs to store the exported sequences.

BASH

cd
mkdir -p analysis/seqs

Run the command to import the raw data located in the directory raw_data and export it to a single QIIME 2 artefact file, combined.qza.

BASH

qiime tools import \
--type 'SampleData[PairedEndSequencesWithQuality]' \
--input-path raw_data \
--input-format CasavaOneEightSingleLanePerSampleDirFmt \
--output-path analysis/seqs/combined.qza

Remove primers


Callout

Remember to ask your sequencing facility if the raw data you get has the primers attached - they may have already been removed.

These sequences still have the primers attached and must be removed prior to denoising. For this workshop, primer trimming was performed using the QIIME 2 cutadapt trim-paired plugin to maintain a fully reproducible workflow within the QIIME 2 framework. Amplicons were generated using standard 16S rRNA gene primers for the v4 region, and the reads returned from the sequencer therefore include these primer sequences at the 5′ ends. Using cutadapt, the specified primer sequence and any bases upstream of the match are removed, with an error rate of 0.10 to balance sensitivity of primer detection with specificity of trimming. Degenerate bases in the primers are accommodated using wildcard matching, and any reads lacking the expected primer sequences are discarded to minimise inclusion of off-target amplification products. A modest 3′ quality trimming threshold (Phred score = 20) is also applied to remove low-quality bases prior to downstream denoising.

It is important to note that these data were generated on an Illumina NextSeq platform, which uses 2-colour chemistry and can produce artificial poly-G tails at the ends of reads under low-signal conditions. The QIIME 2 implementation of cutadapt does not have the –nextseq-trim parameter, which is specifically designed to remove these artefacts. As such, the trimming approach used here represents a simplified, self-contained workflow appropriate for teaching purposes. For production analyses of NextSeq data, best practice is to perform trimming with standalone cutadapt (including –nextseq-trim) prior to importing reads into QIIME 2, as this improves removal of sequencing artefacts and can enhance downstream denoising and taxonomic resolution.

BASH

qiime cutadapt trim-paired \
  --i-demultiplexed-sequences analysis/seqs/combined.qza \
  --p-front-f GTGYCAGCMGCCGCGGTAA \
  --p-front-r GGACTACNVGGGTWTCTAAT \
  --p-error-rate 0.10 \
  --p-overlap 10 \
  --p-match-adapter-wildcards \
  --p-discard-untrimmed \
  --p-quality-cutoff-3end 20 \
  --p-quality-cutoff-5end 0 \
  --p-cores 8 \
  --output-dir analysis/seqs_trimmed \
  --verbose
Caution

The primers specified are the Earth Microbiome Project (EMP) 16S V4 primers (515F (Parada)– 806R (Apprill) targeting the v4 region of the bacterial 16S rRNA gene), which correspond to this specific experiment. Unless you are using these exact primers for your experiment, you need to adapt the code accordingly.

Discussion

The error rate, --p-error-rate, and overlap, --p-overlap, parameters will likely need to be adjusted for your own sample data to maximise the proportion of reads successfully trimmed while avoiding nonspecific matches. Play around with these values and see what happens.

Create and interpret sequence quality data


Create a viewable summary file so the data quality can be checked. Viewing the quality plots generated here helps determine settings for dada2, which we will run next.

Trimmed sequences are processed using the dada2 plugin within QIIME2. dada2 denoises data by modelling and correcting Illumina amplicon sequencing errors, and infers exact amplicon sequence variants (ASVs), resolving differences of as little as a single nucleotide. Its workflow includes filtering, dereplication, paired-end read merging, and reference-free chimera detection, resulting in a feature (ASV) table.

Truncation removes bases from the 3′ end of reads at the specified position. When choosing DADA2 truncation lengths, the first step is to inspect the forward and reverse quality plots you just created. In many modern Illumina datasets, especially after primer and basic quality trimming, these plots may appear relatively flat with consistently high quality across most of the read length. In these cases, there is no obvious “cut point” where quality sharply declines. Instead of looking for a specific quality threshold (for example, Q35), you should choose truncation lengths conservatively, trimming only the very ends of reads if needed while retaining as much high-quality sequence as possible without compromising read overlap.

For paired-end data, an additional and critical consideration is read overlap. After truncation, the forward and reverse reads must still overlap sufficiently to merge. While the absolute minimum overlap is ~12 bp, an overlap of ~50 bp is recommended to ensure robust merging and reduce the risk of losing reads during this step. As a result, truncation lengths are often chosen by balancing two factors: maintaining high-quality sequence and preserving enough overlap for successful merging.

TL;DR: when quality plots are essentially straight lines, truncation is less about identifying where quality drops and more about keeping as much usable sequence as possible while ensuring adequate overlap between reads.

Discussion

Things to look for when choosing truncation lengths

  1. Do the forward and reverse quality profiles show a clear decline near the ends of the reads? If so, truncate before the low-quality tail.
  2. If the quality profiles remain high and relatively flat, avoid trimming too aggressively and retain as much high-quality sequence as possible.
  3. Will the chosen forward and reverse truncation lengths still leave enough overlap for paired-end merging? A minimum overlap of ~50 bp is recommended.

Create a subdirectory in analysis called visualisations to store all files that we will visualise in one place.

BASH

mkdir analysis/visualisations

BASH

qiime demux summarize \
--i-data analysis/seqs_trimmed/trimmed_sequences.qza \
--o-visualization analysis/visualisations/trimmed_sequences.qzv
Challenge

Visualisations: Read quality and demux output

Copy analysis/visualisations/trimmed_sequences.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the trimmed_sequences.qzv file in QIIME 2 View.

Make sure to switch between the “Overview” and “Interactive Quality Plot” tabs in the top left hand corner. Click and drag on the plot to zoom in. Double click to zoom back out to full size. Hover over a box to see the parametric seven-number summary of the quality scores at the corresponding position.

OverviewQualPlotTabs
OverviewQualPlotTabs

Denoising the data


Callout

This step may take a long time to run (i.e. hours), depending on file sizes and available computational power.

Remember to adjust --p-trunc-len-f and --p-trunc-len-r according to your own data.

In the following command, a pooling method of ‘pseudo’ is selected. Pseudo-pooling improves sensitivity to shared low-abundance ASVs across samples while remaining computationally efficient. This is better than the default of ‘independent’ (where samples are denoised independently) when you expect samples in the run to have similar ASVs overall.

Caution

STOP - workshop participants only

Due to time limitations in a workshop setting, please do NOT run the commands below. You will need to access pre-computed files that this command generates by running the following: cd; mkdir analysis/dada2out; cp /mnt/shared_data/pre_computed/dada2out/* analysis/dada2out. If you have accidentally run the command below, ctrl-c will terminate it.

The specified output directory must not pre-exist.

BASH

qiime dada2 denoise-paired \
--i-demultiplexed-seqs analysis/seqs_trimmed/trimmed_sequences.qza \
--p-trunc-len-f xxx \
--p-trunc-len-r xxx \
--p-n-threads 0 \
--p-pooling-method 'pseudo' \
--output-dir analysis/dada2out \
--verbose
Challenge

Calculating truncation lengths

The above code for you likely didn’t work if you didn’t adjust the truncation length from xxx to an integer!

Overlap = (forward truncation length + reverse truncation length) − amplicon length.

For this amplicon, the expected length is ~255 bp. Try a few sensible truncation-length combinations and compare read retention and merging success.

Change this part of the command from:

BASH

--p-trunc-len-f xxx \
--p-trunc-len-r xxx \

to:

BASH

--p-trunc-len-f 210 \ 
--p-trunc-len-r 170 \ 

Tip: Use the up arrow to go back if you tried to run the command before. You can use your arrow keys to edit the command before re-running to replace xxx with the numerical values.

Generate summary files


A metadata file is required which provides the key to gaining biological insight from your data. The file dunnart_metadata.tsv is provided in the home directory of your Nectar instance. This spreadsheet has already been verified using the plugin for Google Sheets, keemei.

Discussion

Things to look for

  1. How many features (ASVs) were generated? Does this seem reasonable for the sample type? High-diversity communities will usually yield more ASVs than low-diversity communities, but very large numbers can also reflect residual noise or non-target amplification.

  2. Do the representative sequences make biological sense? Taxonomic assignments or BLAST hits should broadly match the expected environment or host (for example, marine, soil, gut, or terrestrial communities).

  3. How many reads were retained after filtering, denoising, merging, and chimera removal? If a large proportion of reads were lost (for example, >50%), this may indicate that trimming or truncation settings were too stringent, read quality was poor, or overlap between forward and reverse reads was insufficient.

  4. Did most samples retain enough reads for downstream analysis? Samples with very low final read counts may still be usable in some contexts, but they should be interpreted cautiously and may need to be excluded later.

BASH

qiime metadata tabulate \
--m-input-file analysis/dada2out/denoising_stats.qza \
--o-visualization analysis/visualisations/16s_denoising_stats.qzv \
--verbose
Challenge

Visualisation: Denoising Stats

Copy analysis/visualisations/16s_denoising_stats.qzv to your local computer and view in QIIME 2 View (q2view).


BASH

qiime feature-table summarize \
--i-table analysis/dada2out/table.qza \
--m-metadata-file dunnart_metadata.tsv \
--output-dir analysis/visualisations/summary_table \
--verbose 
Challenge

Visualisations: Feature/ASV summary

Copy analysis/visualisations/summary_table/summary.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the summary.qzv file in QIIME 2 View.

Make sure to switch between the “Overview” and “Feature Detail” tabs in the top left hand corner.
ASV_detailPNG


BASH

qiime feature-table tabulate-seqs \
--i-data analysis/dada2out/representative_sequences.qza \
--o-visualization analysis/visualisations/16s_representative_seqs.qzv \
--verbose
Challenge

Visualisation: Representative Sequences

Copy analysis/visualisations/16s_representative_seqs.qzv to your local computer and view in QIIME 2 View (q2view).

Key Points
  • Confirm input format (CasavaOneEightSingleLanePerSampleDirFmt) when importing.
  • Primer removal is essential; untrimmed primers can disrupt denoising and downstream inference.
  • Inspect interactive quality plots and consider read overlap (~50 bp recommended) when setting truncation lengths for paired‑end merging.
  • Use dada2 denoise‑paired with appropriate pooling (here ‘pseudo’) and monitor read retention in denoising_stats outputs.

Content from Taxonomic Analysis


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • How is taxonomy assigned to the representative sequences?
  • Why must the classifier match the amplified region?
  • What filtering steps are commonly applied to taxonomic results?

Objectives

  • Apply a pre‑trained classifier appropriate for the V4 16S region to assign taxonomy to ASVs.
  • Generate a taxonomic summary visualization and interpret broad community composition.
  • Filter out non‑target classifications (mitochondria, chloroplast) from the feature table.

Assign taxonomy


Here we will classify each identical read or Amplicon Sequence Variant (ASV) to the highest resolution based on a database. Common databases for bacteria datasets are SILVA, Ribosomal Database Project*, or Genome Taxonomy Database. See Porter and Hajibabaei, 2020 for a review of different classifiers for metabarcoding research. The classifier chosen is dependent upon:

  1. Previously published data in a field
  2. The target region of interest
  3. The number of reference sequences for your organism in the database and how recently that database was updated.

A classifier has already been trained for you for the V4 region of the bacterial 16S rRNA gene using the SILVA database. The next step will take a while to run. The output directory cannot previously exist.

n_jobs = 1 This runs the script using all available cores

As of the time of writing, the Ribosomal Data Project website is no longer available. You can find a standalone version of the RDP Classifier 2.14 released in August 2023 on Sourgeforce and Zenodo.

Callout

The classifier used here is only appropriate for the specific 16S rRNA region that this data represents. You will need to train your own classifier for your own data. For more information about training your own classifier, see Extra Information.

BASH

qiime feature-classifier classify-sklearn \
--i-classifier silva_138.2_16s_v4_classifier.qza \
--i-reads analysis/dada2out/representative_sequences.qza \
--p-n-jobs 1 \
--output-dir analysis/taxonomy \
--verbose
Caution

This step often runs out of memory on full datasets. Some options are to change the number of cores you are using (adjust --p-n-jobs) or add --p-reads-per-batch 10000 and try again. The QIIME 2 forum has many threads regarding this issue so always check there was well.

Generate a viewable summary file of the taxonomic assignments.


BASH

qiime metadata tabulate \
--m-input-file analysis/taxonomy/classification.qza \
--o-visualization analysis/visualisations/taxonomy.qzv \
--verbose
Challenge

Visualisation: Taxonomy

Copy analysis/visualisations/taxonomy.qzv to your local computer and view in QIIME 2 View (q2view).

Filtering


Filter out reads classified as mitochondria and chloroplast. Unassigned ASVs are retained. Generate a viewable summary file of the new table to see the effect of filtering.

According to QIIME developer Nicholas Bokulich, low abundance filtering (i.e. removing ASVs containing very few sequences) is not necessary under the ASV model.

BASH

qiime taxa filter-table \
--i-table analysis/dada2out/table.qza \
--i-taxonomy analysis/taxonomy/classification.qza  \
--p-exclude Mitochondria,Chloroplast \
--o-filtered-table analysis/taxonomy/16s_table_filtered.qza \
--verbose

BASH

qiime feature-table summarize \
--i-table analysis/taxonomy/16s_table_filtered.qza \
--m-metadata-file dunnart_metadata.tsv \
--output-dir analysis/visualisations/summary_table_filtered \
--verbose
Challenge

Visualisation: Feature/ASV summary-Filtered

Copy analysis/visualisations/summary_table_filtered/summary.qzv to your local computer and view in QIIME 2 View (q2view).

Key Points
  • Classifier must be trained for the same primer/region used in the dataset; mismatched classifiers give poor results.
  • The classify‑sklearn step can be memory intensive — precomputed classifications are provided for the workshop.
  • Retain unassigned ASVs (they may still be biologically meaningful), but exclude obvious host/plant contaminants.
  • Low‑abundance filtering is generally unnecessary with the ASV paradigm unless specific reasons exist.

Content from Build a phylogenetic tree


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • Why is a phylogenetic tree necessary for some diversity metrics?
  • What are the main steps from sequences to a rooted tree in QIIME 2?

Objectives

  • Align representative sequences with MAFFT, mask non‑informative sites, infer a tree with FastTree, and root the tree.
  • Produce tree artefacts suitable for phylogenetic diversity and UniFrac analyses.

The next step does the following:

  1. Perform an alignment on the representative sequences.
  2. Mask sites in the alignment that are not phylogenetically informative.
  3. Generate a phylogenetic tree.
  4. Apply mid-point rooting to the tree.

A phylogenetic tree is necessary for any analyses that incorporates information on the relative relatedness of community members, by incorporating phylogenetic distances between observed organisms in the computation. This would include any beta-diversity analyses and visualisations from a weighted or unweighted Unifrac distance matrix.

BASH

mkdir analysis/tree

Use one thread only (which is the default action) so that identical results can be produced if rerun.

BASH

qiime phylogeny align-to-tree-mafft-fasttree \
--i-sequences analysis/dada2out/representative_sequences.qza \
--o-alignment analysis/tree/aligned_16s_representative_seqs.qza \
--o-masked-alignment analysis/tree/masked_aligned_16s_representative_seqs.qza \
--o-tree analysis/tree/16s_unrooted_tree.qza \
--o-rooted-tree analysis/tree/16s_rooted_tree.qza \
--p-n-threads 1 \
--verbose
Key Points
  • Phylogenetic distances are required for phylogenetic alpha and beta metrics (Faith’s PD, UniFrac).
  • Use align‑to‑tree‑mafft‑fasttree for a reproducible single‑threaded workflow.
  • Keep track of the unrooted and rooted tree artefacts for downstream use.

Content from Basic visualisations and statistics


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • How can relative abundance barplots aid in exploratory analysis?
  • How do rarefaction and sampling depth choices influence diversity analyses?
  • What statistical tests are appropriate for alpha and beta diversity comparisons here?

Objectives

  • Create interactive taxa barplots and rarefaction curves to assess composition and sequencing depth.
  • Run core‑metrics‑phylogenetic (or equivalent) to compute alpha and beta metrics at a chosen sampling depth.
  • Test for differences in alpha diversity (alpha‑group‑significance) and community composition (beta‑group‑significance / PERMANOVA), and perform differential abundance testing (ANCOM‑BC2).

ASV relative abundance bar charts


Create bar charts to compare the relative abundance of ASVs across samples.

BASH

qiime taxa barplot \
--i-table analysis/taxonomy/16s_table_filtered.qza \
--i-taxonomy analysis/taxonomy/classification.qza \
--m-metadata-file dunnart_metadata.tsv \
--o-visualization analysis/visualisations/barchart.qzv \
--verbose
Challenge

Visualisations: Taxonomy Barplots

Copy analysis/visualisations/barchart.qzv to your local computer and view in QIIME 2 View (q2view). Try selecting different taxonomic levels and metadata-based sample sorting.

Click to view the barchart.qzv file in QIIME 2 View.

Increase the “Bar Width”, select “Captivity” in “Sort Samples By” drop-down menu and explore the resulting barplots by changing the levels in the “Change Taxonomic Level” dropdown menu (Select Level 1, then Level 3, and then Level 5 for example).

barplot1
barplot1

Rarefaction curves

Generate rarefaction curves to determine whether the samples have been sequenced deeply enough to capture all the community members. The max depth setting will depend on the number of sequences in your samples.

Discussion

Things to look for:

  1. Do the curves for each sample plateau? If they don’t, the samples haven’t been sequenced deeply enough to capture the full diversity of the bacterial communities, which is shown on the y-axis.

  2. At what sequencing depth (x-axis) do your curves plateau? This value will be important for downstream analyses, particularly for alpha diversity analyses.

Callout

The value that you provide for –p-max-depth should be determined by reviewing the “Frequency per sample” information presented in the summary.qzv file that was created above after filtering. In general, choosing a value that is somewhere around the median frequency seems to work well, but you may want to increase that value if the lines in the resulting rarefaction plot don’t appear to be levelling out, or decrease that value if you seem to be losing many of your samples due to low total frequencies closer to the minimum sampling depth than the maximum sampling depth.

BASH

qiime diversity alpha-rarefaction \
--i-table analysis/taxonomy/16s_table_filtered.qza \
--i-phylogeny analysis/tree/16s_rooted_tree.qza \
--p-max-depth 200000 \
--p-min-depth 500 \
--p-steps 40 \
--m-metadata-file dunnart_metadata.tsv \
--o-visualization analysis/visualisations/16s_alpha_rarefaction.qzv \
--verbose
Challenge

Visualisation: Rarefaction

Copy analysis/visualisations/16s_alpha_rarefaction.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the 16s_alpha_rarefaction.qzv file in QIIME 2 View.

Select “Animal” in the “Sample Metadata Column” and “observed_features” under “Metric”:

rarefaction
rarefaction

Alpha and beta diversity analysis

The following is taken from the Moving Pictures tutorial and adapted for this data set. QIIME 2’s diversity analyses are available through the q2-diversity plugin, which supports computing alpha- and beta- diversity metrics, applying related statistical tests, and generating interactive visualisations. We’ll first apply the core-metrics-phylogenetic method, which rarefies a FeatureTable[Frequency] to a user-specified depth, computes several alpha- and beta- diversity metrics, and generates principle coordinates analysis (PCoA) plots using Emperor for each of the beta diversity metrics.

The metrics computed by default are:

  • Alpha diversity (operate on a single sample (i.e. within sample diversity)).
    • Shannon’s diversity index (a quantitative measure of community richness)
    • Observed OTUs (a qualitative measure of community richness)
    • Faith’s Phylogenetic Diversity (a qualitative measure of community richness that incorporates phylogenetic relationships between the features)
    • Evenness (or Pielou’s Evenness; a measure of community evenness)
  • Beta diversity (operate on a pair of samples (i.e. between sample diversity)).
    • Jaccard distance (a qualitative measure of community dissimilarity)
    • Bray-Curtis distance (a quantitative measure of community dissimilarity)
    • unweighted UniFrac distance (a qualitative measure of community dissimilarity that incorporates phylogenetic relationships between the features)
    • weighted UniFrac distance (a quantitative measure of community dissimilarity that incorporates phylogenetic relationships between the features)

An important parameter that needs to be provided to this script is --p-sampling-depth, which is the even sampling (i.e. rarefaction) depth that was determined above. As most diversity metrics are sensitive to different sampling depths across different samples, this script will randomly subsample the counts from each sample to the value provided for this parameter. For example, if --p-sampling-depth 500 is provided, this step will subsample the counts in each sample without replacement, so that each sample in the resulting table has a total count of 500. If the total count for any sample(s) are smaller than this value, those samples will be excluded from the diversity analysis. Choosing this value is tricky. We recommend making your choice by reviewing the information presented in the summary.qzv file that was created above. Choose a value that is as high as possible (so more sequences per sample are retained), while excluding as few samples as possible.

BASH

qiime diversity core-metrics-phylogenetic \
  --i-phylogeny analysis/tree/16s_rooted_tree.qza \
  --i-table analysis/taxonomy/16s_table_filtered.qza \
  --p-sampling-depth 100000 \
  --m-metadata-file dunnart_metadata.tsv \
  --output-dir analysis/diversity_metrics

Copy the .qzv files created from the above command into the visualisations subdirectory.

BASH

cp analysis/diversity_metrics/*.qzv analysis/visualisations
Challenge

Visualisations: Unweighted UniFrac Emperor Ordination

To view the differences between sample composition using unweighted UniFrac in ordination space, copy analysis/visualisations/unweighted_unifrac_emperor.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the unweighted_unifrac_emperor.qzv file in QIIME 2 View.

On q2view, select the “Color” tab, choose “Captivity” under the “Select a Color Category” dropdown menu.

unweighted_unifrac_emperor2
unweighted_unifrac_emperor2

Next, we’ll test for associations between categorical metadata columns and alpha diversity data. We’ll do that here for observed ASVs and evenness metrics.

BASH

qiime diversity alpha-group-significance \
  --i-alpha-diversity analysis/diversity_metrics/observed_features_vector.qza \
  --m-metadata-file dunnart_metadata.tsv \
  --o-visualization analysis/visualisations/observed_features-significance.qzv
Challenge

Visualisation: Observed Diversity output

Copy analysis/visualisations/observed_features-significance.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the observed_features-significance.qzv file in QIIME 2 View.

Select “Captivity” under the “Column” dropdown menu.

faith
faith


BASH

qiime diversity alpha-group-significance \
  --i-alpha-diversity analysis/diversity_metrics/evenness_vector.qza \
  --m-metadata-file dunnart_metadata.tsv \
  --o-visualization analysis/visualisations/evenness-group-significance.qzv
Challenge

Visualisation: Evenness output

Copy analysis/visualisations/evenness-group-significance.qzv to your local computer and view in QIIME 2 View (q2view).

Click to view the evenness-group-significance.qzv file in QIIME 2 View.

Select “Captivity” under the “Column” dropdown menu.

evenness
evenness

Next, we’ll analyse sample composition in the context of categorical metadata using a permutational multivariate analysis of variance (PERMANOVA, first described in Anderson (2001)) test using the beta-group-significance command. The following commands will test whether distances between samples within a group are more similar to each other then they are to samples from the other groups. If you call this command with the --p-pairwise parameter, as we’ll do here, it will also perform pairwise tests that will allow you to determine which specific pairs of groups differ from one another, if any. This command can be slow to run, especially when passing --p-pairwise, since it is based on permutation tests. So, unlike the previous commands, we’ll run beta-group-significance on specific columns of metadata that we’re interested in exploring, rather than all metadata columns to which it is applicable. Here we’ll apply this to our unweighted UniFrac distances, using two sample metadata columns, as follows.

BASH

qiime diversity beta-group-significance \
  --i-distance-matrix analysis/diversity_metrics/unweighted_unifrac_distance_matrix.qza \
  --m-metadata-file dunnart_metadata.tsv \
  --m-metadata-column Captivity \
  --o-visualization analysis/visualisations/unweighted-unifrac-captivity-significance.qzv \
  --p-pairwise
Challenge

Visualisation: Captivity significance output and provenance

Copy analysis/visualisations/unweighted-unifrac-captivity-significance.qzv to your local computer and view in QIIME 2 View (q2view).

Finally, we’ll do differential abundance testing with ANCOM-BC2. ANCOM-BC2 is a compositionally-aware linear regression model that allows testing for differentially abundant features across sample groups while also implementing bias correction. This can be accessed using the ancombc2 action in the composition plugin.

We’ll apply ANCOM-BC2 to see which ASV are differentially abundant across Captivity. If you had more than two treatments, you can specify a reference level to define what each group is compared against (--p-reference-levels 'Captivity::Wild'). This is not necessary when you just have two groups.

BASH

qiime composition ancombc2 \
  --i-table analysis/taxonomy/16s_table_filtered.qza \
  --m-metadata-file dunnart_metadata.tsv \
  --p-fixed-effects-formula Captivity \
  --o-ancombc2-output analysis/ancombc2-results.qza

BASH

qiime composition ancombc2-visualizer \
  --i-data analysis/ancombc2-results.qza \
  --i-taxonomy analysis/taxonomy/classification.qza \
  --o-visualization analysis/visualisations/ancombc2-barplot.qzv
Challenge

Visualisation: Differential Abundance Testing

Copy analysis/visualisations/ancombc2-barplot.qzv to your local computer and view in QIIME 2 View (q2view).

Key Points
  • Explore taxonomic composition at multiple taxonomic levels and by metadata categories (e.g. Captivity).
  • Choose sampling/depth parameters informed by the feature‑table summary and rarefaction curves to balance sample retention and depth.
  • Unweighted and weighted UniFrac capture different aspects (presence/absence vs abundance) of community dissimilarity.
  • ANCOM‑BC2 is compositionally aware and suitable for differential abundance while applying bias correction.

Content from Exporting data for further analysis in R


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • Which files are typically exported for downstream R analysis (phyloseq, DESeq2, etc.)?
  • What format conversions are required to obtain usable TSVs and tree files?

Objectives

  • Export the unrooted tree (.nwk), feature table (BIOM), taxonomy and representative sequences to a structured export folder.
  • Convert BIOM to TSV and clean headers so R packages can import tables reliably.
  • Ensure taxonomic rows and ASV columns are consistently ordered for downstream merging.

You need to export your ASV table, taxonomy table, and tree file for analyses in R. Many file formats can be accepted.

Export unrooted tree as .nwk format as required for the R package phyloseq.

BASH

qiime tools export \
  --input-path analysis/tree/16s_unrooted_tree.qza \
  --output-path analysis/export

Create a BIOM table with taxonomy annotations. A FeatureTable[Frequency] artefact will be exported as a BIOM v2.1.0 formatted file.

BASH

qiime tools export \
  --input-path analysis/taxonomy/16s_table_filtered.qza \
  --output-path analysis/export

Then export BIOM to TSV

BASH

biom convert \
-i analysis/export/feature-table.biom \
-o analysis/export/feature-table.tsv \
--to-tsv

Export Taxonomy as TSV

BASH

qiime tools export \
--input-path analysis/taxonomy/classification.qza \
--output-path analysis/export

Delete the header lines of the .tsv files

BASH

sed '1d' analysis/export/taxonomy.tsv > analysis/export/taxonomy_noHeader.tsv
sed '1d' analysis/export/feature-table.tsv > analysis/export/feature-table_noHeader.tsv

Some packages require your data to be in a consistent order, i.e. the order of your ASVs in the taxonomy table rows to be the same order of ASVs in the columns of your ASV table. It’s recommended to clean up your taxonomy file. You can have blank spots where the level of classification was not completely resolved.

Key Points
  • Exported artefacts include: tree (Newick), feature‑table.biom, taxonomy.tsv and representative sequences if required.
  • Use biom convert to produce TSV tables and sed (or equivalent) to remove extra header lines.
  • Confirm consistent ASV order between taxonomy and table files before importing into phyloseq or other R packages.

Content from Extra Information


Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • When and how should you train your own classifier?
  • What special considerations exist for NextSeq data and primer trimming?

Objectives

  • Describe the procedure to extract region‑specific reference reads and train a naive Bayes classifier from SILVA references.
  • Highlight best practices for handling NextSeq 2‑colour chemistry artefacts (e.g. poly‑G tails) and when to perform standalone cutadapt –nextseq‑trim.
Caution

STOP

This section contains information on how to train the classifier for analysing your own data. This will NOT be covered in the workshop.

Train SILVA v138 classifier for 16S/18S rRNA gene marker sequences.


The newest version of the SILVA database (v138) can be trained to classify marker gene sequences originating from the 16S/18S rRNA gene. Reference files silva-138-99-seqs.qza and silva-138-99-tax.qza were downloaded from SILVA and imported to get the artefact files. You can download both these files from here.

Reads for the region of interest are first extracted. You will need to input your forward and reverse primer sequences. See QIIME2 documentation for more information.

BASH

qiime feature-classifier extract-reads \
--i-sequences silva-138-99-seqs.qza \
--p-f-primer FORWARD_PRIMER_SEQUENCE \
--p-r-primer REVERSE_PRIMER_SEQUENCE \
--o-reads silva_138_marker_gene.qza \
--verbose

The classifier is then trained using a naive Bayes algorithm. See QIIME2 documentation for more information.

BASH

qiime feature-classifier fit-classifier-naive-bayes \
--i-reference-reads silva_138_marker_gene.qza \
--i-reference-taxonomy silva-138-99-tax.qza \
--o-classifier silva_138_marker_gene_classifier.qza \
--verbose
Key Points
  • Training a classifier requires region‑matched reference reads and taxonomy; forward/reverse primers must be provided when extracting reads.
  • Classifier training and full classify‑sklearn runs can be computationally demanding — Consider precomputed classifiers or higher memory/CPU resources for large datasets.