A self-guided lab notebook · runs in your browser

Bioinformaticsof gut health.

An eight-module course on the computational pipeline behind gut-health tools like GMWI2 — from a raw sequencer read to a trained, cross-validated classifier. Parse real sequence data with Biopython, hand-build a presence/absence matrix, train your own Lasso logistic regression with scikit-learn, and see exactly how a 2024 published tool (Mayo Clinic's Gut Microbiome Wellness Index 2) does the same thing at real scale. Runs natively in-browser; no local installation required.

Start the course
Fig. 1 — The pipeline this course builds, stage by stage
01Raw reads
02QC + taxonomy
03Presence /
absence
04Health-species score
05Trained classifier
06GMWI2,
for real
Each stage above is one or more notebooks in this course — by the time you reach Notebook 06, you've already hand-built every step that a real run of GMWI2 performs automatically.

Specimen log — course order

№ 00 Welcome
How this course works, and how it differs from the biology course
№ 01 Reading Raw Sequence Data
FASTA/FASTQ, Biopython, what a sequencing read actually is
hands-on
№ 02 From Reads to a Taxonomy Table
QC, MetaPhlAn3, why GMWI2 needs shotgun data, not 16S
conceptual
№ 03 From Abundance to Presence/Absence
Binarizing a species table, by hand, on real-shaped synthetic data
hands-on
№ 04 Health- vs. Disease-Associated Species
A simplified log-ratio health score, hand-built
hands-on
№ 05 Training a Health Classifier
Lasso logistic regression, cross-validation, balanced accuracy, with scikit-learn
hands-on
№ 06 Running GMWI2 for Real
The real CLI, real output files, real 2024 Nature Communications citation
conceptual
№ 07 Limits, Critiques, and What's Next
What GMWI2 can't tell you, and where health-index research is headed
conceptual
Nothing to install. Every hands-on notebook here uses only pandas, numpy, scikit-learn, and Biopython — all bundled in the in-browser Python runtime already. No setup cells, no waiting on installs.