Extra Information
Last updated on 2026-10-06 | Edit this page
Estimated time: 12 minutes
Overview
Questions
- When and how should you train your own classifier?
- What special considerations exist for NextSeq data and primer trimming?
Objectives
- Describe the procedure to extract region‑specific reference reads and train a naive Bayes classifier from SILVA references.
- Highlight best practices for handling NextSeq 2‑colour chemistry artefacts (e.g. poly‑G tails) and when to perform standalone cutadapt –nextseq‑trim.
STOP
This section contains information on how to train the classifier for analysing your own data. This will NOT be covered in the workshop.
Train SILVA v138 classifier for 16S/18S rRNA gene marker sequences.
The newest version of the SILVA database (v138) can be
trained to classify marker gene sequences originating from the 16S/18S
rRNA gene. Reference files silva-138-99-seqs.qza and
silva-138-99-tax.qza were downloaded from
SILVA and imported to get the artefact files. You can download both
these files from here.
Reads for the region of interest are first extracted. You will need to input your forward and reverse primer sequences. See QIIME2 documentation for more information.
BASH
qiime feature-classifier extract-reads \
--i-sequences silva-138-99-seqs.qza \
--p-f-primer FORWARD_PRIMER_SEQUENCE \
--p-r-primer REVERSE_PRIMER_SEQUENCE \
--o-reads silva_138_marker_gene.qza \
--verbose
The classifier is then trained using a naive Bayes algorithm. See QIIME2 documentation for more information.
BASH
qiime feature-classifier fit-classifier-naive-bayes \
--i-reference-reads silva_138_marker_gene.qza \
--i-reference-taxonomy silva-138-99-tax.qza \
--o-classifier silva_138_marker_gene_classifier.qza \
--verbose
- Training a classifier requires region‑matched reference reads and taxonomy; forward/reverse primers must be provided when extracting reads.
- Classifier training and full classify‑sklearn runs can be computationally demanding — Consider precomputed classifiers or higher memory/CPU resources for large datasets.