Extra Information

Last updated on 2026-10-06 | Edit this page

Overview

Questions

  • When and how should you train your own classifier?
  • What special considerations exist for NextSeq data and primer trimming?

Objectives

  • Describe the procedure to extract region‑specific reference reads and train a naive Bayes classifier from SILVA references.
  • Highlight best practices for handling NextSeq 2‑colour chemistry artefacts (e.g. poly‑G tails) and when to perform standalone cutadapt –nextseq‑trim.
Caution

STOP

This section contains information on how to train the classifier for analysing your own data. This will NOT be covered in the workshop.

Train SILVA v138 classifier for 16S/18S rRNA gene marker sequences.


The newest version of the SILVA database (v138) can be trained to classify marker gene sequences originating from the 16S/18S rRNA gene. Reference files silva-138-99-seqs.qza and silva-138-99-tax.qza were downloaded from SILVA and imported to get the artefact files. You can download both these files from here.

Reads for the region of interest are first extracted. You will need to input your forward and reverse primer sequences. See QIIME2 documentation for more information.

BASH

qiime feature-classifier extract-reads \
--i-sequences silva-138-99-seqs.qza \
--p-f-primer FORWARD_PRIMER_SEQUENCE \
--p-r-primer REVERSE_PRIMER_SEQUENCE \
--o-reads silva_138_marker_gene.qza \
--verbose

The classifier is then trained using a naive Bayes algorithm. See QIIME2 documentation for more information.

BASH

qiime feature-classifier fit-classifier-naive-bayes \
--i-reference-reads silva_138_marker_gene.qza \
--i-reference-taxonomy silva-138-99-tax.qza \
--o-classifier silva_138_marker_gene_classifier.qza \
--verbose
Key Points
  • Training a classifier requires region‑matched reference reads and taxonomy; forward/reverse primers must be provided when extracting reads.
  • Classifier training and full classify‑sklearn runs can be computationally demanding — Consider precomputed classifiers or higher memory/CPU resources for large datasets.