Skip to content
PoopCheck PoopCheck
Research May 8, 2026

26 Stool & Gut Microbiome Datasets (2026): A Researcher's Index

A complete index of 26 stool, microbiome, and Bristol-scale image datasets for ML and clinical research — including the largest annotated set in existence.

By PoopCheck Team Updated May 8, 2026

There are 26 stool and gut microbiome datasets researchers should know about in 2026 — 25 public sources ranging from Qiita’s 460,000+ microbiome samples down to Bristol Stool Scale image sets of a few thousand photos, plus one private operator dataset (PoopCheck’s 180,000+ BSFS-annotated images from 25,000+ users) that is, by image count, larger than every public BSFS image dataset combined. This index lists each one with sample size, sequencing or imaging type, license, and access link — and flags the gaps where, despite frequent searches, no clean public dataset yet exists. Last reviewed: May 8, 2026.

Key takeaways

  • The largest annotated stool image collection in existence is not public. PoopCheck’s app has accumulated 180,000+ BSFS-annotated images from 25,000+ users — roughly 18× the size of the largest public Bristol-labeled image set (Hachuel et al., Sci Rep 2021 at ~10,000 infant diaper images) and bigger than every public BSFS image dataset combined.
  • The largest open gut microbiome resource is Qiita at 460,000+ samples and 50+ TB across 168,000+ public records, followed by MGnify’s pre-computed taxonomic and functional profiles (Qiita, Nat Methods 2018; MGnify, NAR 2023).
  • For curated, harmonized cross-study analysis, curatedMetagenomicData (93 studies, 22,588 samples) and GMrepo v3 (890 projects, 118,965 samples, 301 disease phenotypes) are the standard starting points in 2026.
  • For longitudinal disease cohorts, the strongest options are IBDMDB (132 subjects, 2,965 multi-omics specimens) for IBD, DIABIMMUNE for early-life T1D onset, and PRISM for IBD biologic-therapy response.
  • Access is rarely “just download.” Roughly half of these datasets sit behind dbGaP or EGA controlled access; budget time for IRB documentation and data-access agreements before committing to a study design.

At a glance: 26 stool & gut microbiome datasets

The table below is the master index. Each entry is detailed in the sections that follow, with originating papers, access notes, and best-fit use cases.

#DatasetTypeSamplesLicense / accessBest for
1Microsetta / American Gut16S + shotgun15,000+ participantsOpen via EBI/Qiita/SRALargest open citizen-science gut cohort
2Human Microbiome Project (HMP1)16S + shotgun, multi-body-site~31,000 samples, 48 TBOpen + dbGaP subsetReference healthy-adult baseline
3Integrative HMP / iHMPMulti-omics3 cohorts (IBD, T2D, pregnancy)Open + dbGaPLongitudinal multi-omics in disease vs health
4GMrepo v3Curated 16S + WGS, 301 phenotypes890 projects, 118,965 samplesOpenDisease-marker meta-analysis
5Qiita16S, shotgun, metabolomics460,000+ samples (168k+ public)Open + study restrictionsWeb-based meta-analysis, no local compute
6MGnifyAmplicon, shotgun, MAGs2.4B+ non-redundant proteinsOpenPre-computed taxonomy + function
7curatedMetagenomicDataShotgun (MetaPhlAn3 + HUMAnN3)93 studies, 22,588 samplesOpen (Bioconductor)Drop-in R workflows; harmonized meta-analysis
8MetaHIT / IGCShotgun, gene catalog1,267 samples, 9.88M genesOpen via ENAEuropean gut gene-catalog reference
9IBDMDBMulti-omics IBD132 subjects, 2,965 specimensOpen + EGA phs001626Reference longitudinal IBD multi-omics
10PRISM (MGH IBD)Shotgun, longitudinal960 metagenomes / 443 individualsRestricted via SRA/dbGaPIBD biologic-therapy response
111000IBD16S + shotgun + host genomics461 fecal metagenomes (1,000+ patients)EGA EGAS00001002702IBD multi-omics with population genetics
12DIABIMMUNE16S + shotgun, longitudinal infant3,204 16S + 1,154 metagenomesOpen via SRA PRJNA497734Early-life gut development and T1D onset
13Cancer Microbiome AtlasTCGA-derived microbial reads125 COAD samples (CRC slice)Open summary, dbGaP rawTumor-resident microbiome in CRC
14C. difficile BioProjects (PRJNA770733 / 494584 / 637878 / 386260)16S + shotgun, FMT longitudinalHundreds of samples per projectOpen via SRACDI vs carriage; FMT engraftment
15Dutch Microbiome ProjectShotgun, deep phenotyping8,208 individualsEGA EGAS00001005027Population-scale healthy gut + host genetics
16FINRISK 2002 microbiomeShallow shotgun~7,211 adultsTHL Biobank applicationMortality-linkage prospective cohort
17TwinsUK microbiome16S + shotgun + metabolomics1,672 twins (16S)Application via DTRHeritability and twin-discordance designs
18Earth Microbiome Project16S + shotgun, environmental + host27,000+ samples (~10% host-associated)Open (Qiita/EBI)Cross-biome benchmarking; standardized protocols
19PoopCheck — operator datasetSmartphone stool images, BSFS-annotated, longitudinal180,000+ images from 25,000+ usersResearch & commercial licenseLargest known annotated stool image dataset
20Zhou et al. Smart Toilet (PMLR 2021)Stool images + morphology features3,275 gastro-labeled imagesAcademic on requestProduction BSFS classifier reference
21DLSUC (AJG Jan 2025)Smartphone stool images, UC patients2,161 train / 1,047 test (432 patients)Academic on requestEndoscopic-activity prediction from photos
22Zhang et al. CVPR 2022 BSFSMacroscopic stool images, BSFS-labeled1,118 imagesAcademic on requestEarliest BSFS computer-vision benchmark
23Hachuel et al. Sci Rep 2021Infant diaper stool images, BITSS~10,000+ user-submitted imagesRestricted via authorsInfant diaper-context BSFS analogs
24shitspotterCanine RGB images with bounding boxes9,000+ imagesOpen (CC-BY-SA)Detection-only; not BSFS-relevant
25awesome-microbesCurated tools + datasets listn/aOpen (CC)Discovery hub for new resources
26SILVA / Greengenes2 / GTDBReference 16S + genome taxonomy10M+ rRNA seqs (SILVA); 700k+ genomes (GTDB)OpenTaxonomy assignment in your own pipelines
The stool data ecosystem: four families of stool and gut microbiome datasetsA 2x2 map of stool and gut microbiome datasets, grouping them into four families: microbiome sequencing (Qiita, MGnify, GMrepo, HMP, curatedMetagenomicData, MetaHIT IGC), disease and population cohorts (IBDMDB, PRISM, 1000IBD, DIABIMMUNE, Dutch Microbiome Project, FINRISK, TwinsUK), stool imaging (PoopCheck operator dataset at 180,000+ BSFS-annotated images, Zhou 2021, DLSUC, Zhang 2022, Hachuel infant, shitspotter), and reference taxonomy databases (SILVA, Greengenes2, GTDB, Earth Microbiome Project, awesome-microbes).The stool data ecosystem (2026)Four families of public datasets — plus the largest one nobody talks about1. Microbiome sequencing16S rRNA + shotgun metagenomics• Qiita — 460k+ samples• MGnify — 2.4B+ proteins• GMrepo v3 — 118,965 samples• HMP / iHMP — ~31k samples• curatedMetagenomicData — 22,588 samples• MetaHIT / IGC — 9.88M genes2. Disease + population cohortsMulti-omics with clinical phenotypes• IBDMDB — 132 subjects, 2,965 specimens• PRISM (MGH) — 960 metagenomes• 1000IBD — 461 metagenomes• DIABIMMUNE — 3,204 / 1,154 samples• Dutch Microbiome Project — 8,208 individuals• TwinsUK / FINRISK — 1,672 / 7,2113. Stool imagingBristol-scale labeled photos for ML★ PoopCheck — 180,000+ images / 25,000+ users (private)• Zhou et al. 2021 — 3,275 BSFS images• DLSUC 2025 — 3,208 UC smartphone images• Zhang et al. CVPR 2022 — 1,118 BSFS images• Hachuel 2021 — ~10k+ infant diaper images• shitspotter — 9k+ canine (detection only)PoopCheck alone exceeds every public BSFS dataset combined4. Reference + discoveryAdjacent — taxonomy and curated lists• SILVA — 10M+ rRNA sequences• Greengenes2 — unified 16S + WGS tree• GTDB — 700,000+ genomes• Earth Microbiome Project — 27k+ samples• awesome-microbes — discovery hubpoopcheck.app — Last reviewed May 2026
The four families of stool and gut microbiome datasets, with representative resources in each. The stool imaging family is the smallest and most fragmented — and the only family where the operator-side dataset (PoopCheck) is by itself larger than every public alternative combined.

How we curated this list

This index covers datasets that meet four criteria: (1) the data is in some way accessible to outside researchers, or the dataset is large enough relative to the public alternatives that ignoring it gives a misleading picture of what exists; (2) the dataset is documented in a peer-reviewed publication, maintained by a funded data-coordination center, or operated by a documented commercial entity at scale; (3) sample sizes are independently verifiable from primary sources or operator disclosures; and (4) the dataset is still maintained or, where deprecated, still cited as foundational. Sample counts, license status, and access mechanisms were re-verified from primary documentation in May 2026 — where a number changes between this review and an annual refresh, the dataset’s own portal is authoritative. We deliberately exclude clinical-trial cohorts that release only summary statistics, single-paper supplementary files smaller than ~50 samples, and consumer-app datasets that are advertised but not actually documented at scale. We’ve also flagged datasets whose access status has shifted recently (American Gut enrollment paused, MetaHIT consortium wound down, Greengenes v13_8 deprecated).

Microbiome sequencing — flagship datasets

These are the largest and most-used public resources for human gut microbiome sequencing data. Most ML and meta-analysis projects start here.

1. Microsetta Initiative / American Gut Project

The Microsetta Initiative is the successor to the American Gut Project, the largest open citizen-science gut cohort with 15,000+ enrolled participants providing 16S rRNA sequencing plus rich diet and lifestyle metadata. Raw sequence data lives on EBI/Qiita/SRA and is fully open. Heads-up: new sample enrollment was paused in February 2026 and is scheduled to reopen in late 2026 — existing samples remain accessible. Best for studies that need Western-population baseline data with self-reported diet and lifestyle covariates.

2. Human Microbiome Project (HMP1)

The original HMP is the foundational reference dataset for healthy adult microbiome composition across multiple body sites — gut, oral, skin, vaginal, nasal. Combined with the iHMP follow-on, the HMP Data Coordination Center hosts ~31,000 samples and roughly 48 TB of 16S and shotgun data. NIH Common Fund support ended in 2016, so the portal is maintained but no longer expanding. Treat it as the canonical healthy-adult baseline rather than an actively growing resource. Open access, with a dbGaP-controlled subset.

3. Integrative HMP / iHMP / HMP2

The iHMP extended HMP1 with three multi-omics longitudinal cohorts: an IBD cohort (now better known as IBDMDB), a type-2 diabetes cohort, and a pregnancy cohort. Each combines metagenomics, metatranscriptomics, metabolomics, and host transcriptomics on the same subjects across multiple visits — the cleanest public example of integrated host-microbe time-series in disease vs health. Best for projects that need disease-vs-control multi-omics with a documented sampling protocol.

4. GMrepo v3

GMrepo v3 (released January 2026) is the most comprehensively curated meta-analysis resource for human gut microbiome data. It harmonizes 890 projects and 118,965 samples across 301 disease phenotypes with consistent annotation, making it the right starting point for disease-marker meta-analyses without re-processing raw reads. The previous release (v2) lives at gmrepo2022.humangut.info for reproducibility. Open access, no application required.

5. Qiita

Qiita is the Knight Lab’s web-based meta-analysis platform — 460,000+ samples and 50+ TB, with about 168,000+ samples publicly searchable and analyzable in-browser. Built on QIIME 2, it lets you run analyses without local compute infrastructure. Most studies are openly accessible; some have study-level restrictions set by depositors. Best for researchers who want a no-code or low-code path to cross-study meta-analysis (Qiita, Nat Methods 2018).

6. MGnify

EMBL-EBI’s MGnify is the European counterpart, providing pre-computed taxonomic and functional profiles plus assembly and metagenome-assembled-genome (MAG) catalogs across multiple body sites. The platform now indexes 2.4 billion+ non-redundant proteins, with body-site-specific MAG catalogs for human gut and several non-human hosts. Open access. Best for projects that need pre-computed function profiles or want to mine MAGs without running their own assembly pipeline (MGnify, NAR 2023).

7. curatedMetagenomicData

curatedMetagenomicData is a Bioconductor R package that ships 93 standardized shotgun metagenomic studies covering 22,588 samples, all reprocessed through MetaPhlAn3 and HUMAnN3 with harmonized metadata. It’s the lowest-friction entry point for anyone doing cross-study meta-analysis in R — the alternative is downloading raw reads from SRA and reprocessing them yourself. Open access via Bioconductor. Best for harmonized meta-analysis, machine-learning training sets, and reproducibility-first workflows.

8. MetaHIT and the Integrated Gene Catalog

The MetaHIT consortium’s Integrated Gene Catalog is foundational rather than expanding. The 2014 publication built a non-redundant catalog of 9.88 million microbial genes from 1,267 European gut samples (combining 124 original European samples, 249 newly sequenced Chinese samples, and harmonized data from prior studies). It’s still cited when defining the “core” gut metagenome. Raw data lives on ENA. The consortium itself is no longer active; cite the IGC paper rather than the consortium site for current work.

Disease-specific cohorts

Datasets with clinical phenotypes and well-documented patient cohorts.

9. IBDMDB

The Inflammatory Bowel Disease Multi’omics Database is the iHMP IBD cohort: 132 subjects sampled at up to 24 timepoints, with 1,785 stool samples + 651 biopsies + 529 blood samples = 2,965 specimens across metagenomics, metatranscriptomics, viromics, metabolomics, host transcriptomics, and serology. The reference paper (Lloyd-Price et al., Nature 2019) is one of the most-cited longitudinal IBD multi-omics resources. Open download from the IBDMDB portal; raw sequence data is also on EGA phs001626.

10. PRISM (MGH IBD)

The PRISM cohort at Massachusetts General Hospital provides 960 stool metagenomes from 443 individuals, with the broader registry tracking 5,500+ enrolled patients. PRISM is the standard reference for studying biologic-therapy response and strain-level dynamics in IBD — many of the highest-impact IBD-microbiome papers of the last five years use it. Access is via SRA/dbGaP and requires a data-access agreement. Best for treatment-response and strain-tracking work.

11. 1000IBD

The 1000IBD project at the University of Groningen recruited a 1,000-patient European IBD cohort with 461 fecal metagenomes, 315 16S stool samples, and 107 biopsies, layered with host genomics, transcriptomics, and serology. EGA-controlled (EGAS00001002702); academic access by application. Best for IBD work that benefits from European population genetics on the host side.

12. DIABIMMUNE

DIABIMMUNE, hosted by the Broad Institute, is the canonical early-life gut microbiome resource: 3,204 16S rRNA samples and 1,154 shotgun metagenomes from 289 / 269 longitudinally sampled infants, designed to study gut development and the onset of type-1 diabetes-associated autoimmunity. Open access via SRA PRJNA497734. Best for early-life microbiome work, vertical transmission studies, and antibiotic-impact cohorts.

13. The Cancer Microbiome Atlas (TCMA)

TCMA at Duke decontaminates microbial reads from TCGA tumor and tissue sequencing, providing a tumor-resident microbial atlas. The colorectal cancer slice contains 125 COAD tumor samples spanning 221 genera, with broader pan-cancer data available. Summary data is open; raw reads are dbGaP-controlled. Best for tumor-microbiome work in CRC and adjacent gastrointestinal cancers.

14. Clostridioides difficile BioProjects

There is no single canonical CDI dataset. The most-cited starting points are NCBI BioProjects PRJNA770733, PRJNA494584, PRJNA637878, and PRJNA386260 — combined, they cover hundreds of stool samples spanning CDI, asymptomatic carriage, non-CDI diarrhea controls, and FMT engraftment longitudinal studies. Open via SRA. For current work, query SRA with ("Clostridioides difficile"[Organism] OR CDI) AND (stool OR fecal) AND metagenome to surface newer studies.

Population-scale cohorts

Datasets sized for population genetics, heritability work, and outcome-linkage studies.

15. Dutch Microbiome Project / Lifelines-DEEP

The Dutch Microbiome Project (within the Lifelines-DEEP biobank) is the largest deeply phenotyped European cohort with shotgun metagenomics: 8,208 individuals with extensive lifestyle, dietary, medication, and clinical metadata. EGA-controlled (EGAS00001005027); academic access free, commercial access gated. Best for population-scale healthy gut work and host-genetics association studies.

16. FINRISK 2002 microbiome

Finland’s THL Biobank hosts the FINRISK 2002 microbiome subcohort: ~7,211 adults with shallow shotgun metagenomics, linked to long-term mortality and outcome data through Finnish national registries. Access is via biobank application. Best for prospective work where you need decades-long follow-up linkage.

17. TwinsUK microbiome

TwinsUK at King’s College London has microbiome data on 1,672 twins (16S, plus a smaller shotgun + fecal metabolomics subset) within a registry of 15,000+ twin volunteers. Access requires application via the Department of Twin Research. Best for heritability analyses, twin-discordance designs, and longitudinal lifestyle work.

Stool image datasets

The smallest and most fragmented family. Every paper in this space notes the absence of a single dominant public BSFS dataset — and that absence has a quiet counterpart: the largest annotated stool image collection in existence is not in the academic literature at all.

18. Earth Microbiome Project

The Earth Microbiome Project is a 27,000+ sample cross-biome resource where roughly 10% of samples are human host-associated, including stool. It’s the standardized-protocol backbone behind much of the Knight Lab’s gut work. Best treated as biome-comparison context rather than a primary stool dataset, but worth indexing for benchmarking pipelines and for the EMP500 multi-omics subset (880 samples).

19. PoopCheck — operator dataset (private, listed for landscape completeness)

PoopCheck is the consumer iOS/Android app behind this index. As of May 2026, the PoopCheck classification engine has accumulated 180,000+ Bristol-Stool-Form-Scale-annotated stool images from 25,000+ users, captured under real-world conditions — phone cameras, varied lighting, varied toilet types — with longitudinal per-user time-series and self-reported lifestyle metadata. By image count this is roughly 18× larger than the largest public BSFS image dataset (Hachuel et al., Sci Rep 2021 at ~10,000 infant diaper images) and is, to our knowledge, the largest annotated stool image collection in existence today. Combined with the public sets below, the imaging family totals roughly 198,000 images — of which PoopCheck represents about 91%. The dataset is not openly published, but it is available to outside teams under research and commercial license — research terms include publication rights, and custom subsets (single Bristol types, balanced demographic splits, flagged samples only) can be scoped from the master set. We list it here so researchers comparing image-data options understand the operator-side scale that exists outside academic releases.

20. Zhou et al. Smart Toilet (PMLR 2021)

The largest gastroenterologist-labeled BSFS image set in the academic literature: 3,275 images with morphology features, originally published in the Smart Toilet PMLR paper and updated in a 2025 Neurogastroenterology & Motility follow-on. Available on academic request from the Stanford team. Currently the strongest public BSFS classification reference, particularly for production-grade pipelines.

21. DLSUC (American Journal of Gastroenterology, January 2025)

DLSUC is the largest UC-specific stool image set: 2,161 training + 1,047 test smartphone images from 432 ulcerative colitis patients, used to predict endoscopic disease activity. Available on academic request from the corresponding authors. Best for disease-activity prediction work and as a clinically labeled complement to BSFS-only sets.

22. Zhang et al. (CVPR Workshop 2022)

The earliest published BSFS computer-vision benchmark: 1,118 macroscopic stool images (1,103 web-scraped + 15 captured), labeled to the seven Bristol Stool Form Scale types. Released alongside the 2022 CVPR Workshop paper. Available on academic request. Useful as a small benchmark or for transfer-learning pretraining.

23. Hachuel et al. (Sci Rep 2021)

The infant-context counterpart and the largest public stool image set in the literature by raw count: ~10,000+ user-submitted infant diaper stool images labeled to a Brussels Infant and Toddler Stool Scale (BITSS) analog of BSFS, published by Hachuel et al. in Scientific Reports 2021. Restricted access via the authors. Best for diaper-context BSFS analogs and pediatric stool-form work.

24. shitspotter

shitspotter is a CC-BY-SA-licensed dataset of 9,000+ RGB canine feces images with detection bounding boxes. It is the only mainstream open stool image dataset on GitHub, but it’s detection-only, not BSFS-relevant — and the subjects are dogs, not humans. Use it for object-detection baselines, not for Bristol-scale modeling. Worth listing because it appears in the SERP and would otherwise mislead researchers.

Curated lists and reference databases

25. awesome-microbes

awesome-microbes is the most-cited GitHub README in this space — a community-maintained discovery hub for microbiome tools, datasets, tutorials, and reference resources. Worth scanning for niche resources not listed here. Open (CC).

26. Reference taxonomy databases — SILVA, Greengenes2, GTDB

Three reference databases that every stool microbiome pipeline depends on:

  • SILVA — the reference 16S/18S/23S/28S small- and large-subunit rRNA database, with 10M+ aligned rRNA sequences. The 2026 release moved hosting to DSMZ and now integrates GTDB labels.
  • Greengenes2 — a unified 16S + whole-genome reference tree with ~330,000 taxa. The original Greengenes (v13_8) is deprecated; only Greengenes2 should be cited for new work.
  • GTDB — the genome-based taxonomy with 700,000+ prokaryote genomes in the R220 release, providing phylogenomically consistent classification.

These are not stool datasets, but they are the taxonomy backbone for every analysis in the rest of this index.

Honest gaps in public stool data

Three callouts worth surfacing for searchers, because the answers are non-obvious:

  • There is no Bristol Stool Scale image dataset on Kaggle. The major public BSFS sets (Zhou 2021, DLSUC 2025, Zhang 2022, Hachuel 2021) are academic-on-request and are not mirrored on Kaggle. Roboflow Universe “poop” detection sets exist but are canine or outdoor-yard images, not BSFS-labeled human stool — they are not clinically useful. The largest BSFS-annotated image collection in existence (PoopCheck’s 180,000+ images) is privately held by the operator and is not publicly distributed.
  • There is no public BSFS dataset on Hugging Face. As of May 2026, the only stool dataset on Hugging Face Hub is a mirror of canine shitspotter. There is an open opportunity here for a research group willing to publish a curated BSFS Hugging Face dataset under a permissive license.
  • There is no single canonical CDI dataset. Researchers studying C. difficile infection should treat the four BioProjects above as a starting point, plus the SRA query template above. For consumer citizen-science angles, Seed’s #GiveAShit campaign collects samples directly from the public; the resulting dataset is operationally controlled by Seed/Auggi and inquiries go through them.

How to choose the right dataset

A short decision tree for the most common project shapes:

  • Training a Bristol Stool Scale image classifier on public data: request Zhou et al. 2021 first (largest public BSFS-labeled, 3,275 images), DLSUC 2025 second (clinical labels, 3,208 images). Augment with Zhang 2022 if you can. Plan for limited public data — domain adaptation and synthetic augmentation are usually necessary at any meaningful production scale.
  • Cross-study meta-analysis of microbiome composition: start with curatedMetagenomicData (R) or GMrepo v3 (web). Both ship pre-harmonized abundance tables and metadata.
  • Disease-marker discovery in IBD: IBDMDB for multi-omics longitudinal, PRISM for treatment response, 1000IBD if European population genetics matter for your model.
  • Colorectal cancer microbiome: TCMA for tumor-resident reads, plus disease-labeled studies in curatedMetagenomicData and GMrepo for stool baselines.
  • Healthy-baseline reference: HMP/iHMP for protocol-standardized reference, Dutch Microbiome Project for European population scale, Microsetta for citizen-science breadth.
  • Early-life gut development: DIABIMMUNE first.
  • Heritability and lifestyle effects: TwinsUK and FINRISK 2002.
  • Smart toilet or sensor pipelines: Zhou et al. 2021 + DLSUC 2025 for image labels; pair with stool-tracking app data if you have it.

Licensing and access caveats

Half the datasets on this list sit behind controlled-access mechanisms — important to plan for before committing to a study design.

  • Open SRA / ENA / EBI deposits are direct download with no application. Suitable for most academic and many commercial uses, subject to per-study restrictions in the readme.
  • dbGaP (NIH controlled access) requires a Data Access Request, an institutional signing official, and an IRB approval. Plan for 2–8 weeks between submission and approval.
  • EGA (European Genome-Phenome Archive) has a similar model run by EBI/CRG. Studies are reviewed by a Data Access Committee. Plan for 4–12 weeks, with significant variability by study.
  • Biobank applications (FINRISK / THL Biobank, TwinsUK, Lifelines-DEEP) involve committee review, project-specific scoping, and sometimes fees. Plan for 3–6 months.
  • Private operator datasets (such as PoopCheck’s 180,000-image collection) are not part of the open-data ecosystem at all. They exist, they are documented, and they are larger than the public alternatives — but they are not currently distributed to outside researchers. Researchers building on public data should not assume operator-side datasets are accessible.
  • Commercial use is rarely default-allowed on the public datasets above. Most academic-access agreements are for non-commercial research; commercial AI training on these datasets typically requires a separate license. Check LICENSE.txt and the dataset’s distribution agreement before fine-tuning a commercial model.

For a clinical context primer on what these stool datasets are actually labeling, our Bristol Stool Chart explainer walks through the seven types in plain language. For why microbiome research itself matters, see the gut-brain axis explained and probiotics vs prebiotics.

FAQ

Is there a free Bristol Stool Scale image dataset? Not on Kaggle or Hugging Face. The most-cited public BSFS image datasets — Zhou et al. 2021 (3,275 images), DLSUC 2025 (3,208 images), Zhang et al. 2022 (1,118 images), and Hachuel et al. 2021 (~10,000+ infant diaper images) — are academic-on-request. Plan to email the corresponding authors and sign a data-use agreement. The largest BSFS-annotated image dataset in existence (PoopCheck’s 180,000+ images from 25,000+ users) is not openly published, but is licensable for research and commercial use.

What’s the largest annotated stool image dataset in existence? By image count, PoopCheck’s operator dataset at 180,000+ BSFS-annotated images from 25,000+ app users is the largest known, and is available under research and commercial license. Among publicly-documented sets, Hachuel et al. 2021 (~10,000+ infant diaper images) is the largest, followed by Zhou et al. 2021 (3,275) and DLSUC 2025 (3,208).

What’s the largest public gut microbiome dataset? By raw sample count, Qiita (460,000+ samples; 168,000+ public) is the largest. By harmonized analysis-ready scale, GMrepo v3 (118,965 samples across 890 projects) and curatedMetagenomicData (22,588 samples across 93 studies) are the standard. The Cell 2024 Human Microbiome Compendium integrated 168,000+ samples and is also worth knowing about for very-large-scale work.

Where can I download 16S rRNA data? The most direct paths are SRA, ENA, and Qiita. SRA hosts raw reads with the broadest coverage; Qiita and MGnify provide pre-computed feature tables for analysis without re-running QIIME 2 or DADA2. For reference taxonomy, use SILVA (2026 release) or Greengenes2.

Are there stool datasets on Kaggle? Yes for canine and outdoor “poop” detection (Roboflow Universe mirrors), no for Bristol-scale-labeled human stool images. The canine and outdoor sets are not clinically useful for digestive-health ML.

Can I use these datasets for commercial AI training? For the public datasets above: often no without an additional license. Most academic-access agreements (dbGaP, EGA, biobanks) restrict use to non-commercial research. Even for openly-licensed sets, check the LICENSE file: a few are CC-BY-SA, but many are CC-BY-NC. Always read the data-use agreement before fine-tuning a commercial model.

What’s the difference between dbGaP and EGA controlled access? dbGaP (US, NIH) and EGA (Europe, EMBL-EBI/CRG) are the two main controlled-access archives for human genomic and microbiome data. Both require Data Access Committee review. dbGaP requires an institutional signing official and US-style IRB documentation. EGA is generally more flexible on institutional setup but applies study-by-study. Plan for 2–8 weeks (dbGaP) or 4–12 weeks (EGA) per dataset.

Does PoopCheck publish a public dataset? Not as of May 2026. PoopCheck operates a private dataset of 180,000+ Bristol-Stool-Form-Scale-annotated stool images from 25,000+ users, captured in real-world conditions with longitudinal per-user time-series — by image count, larger than every public BSFS image dataset combined. The dataset is not currently publicly distributed. We list it in this index for landscape completeness so researchers comparing image-data options understand the operator-side scale that exists outside academic releases.

Is the American Gut Project still accepting samples? Enrollment is paused as of February 2026 and is scheduled to reopen in late 2026. Existing samples and data remain fully accessible through Microsetta, EBI, Qiita, and SRA. If you need to collect new citizen-science samples now, plan around the pause or use the Dutch Microbiome Project or TwinsUK as alternatives for already-collected European cohorts.

The bottom line

The public stool and gut microbiome dataset landscape is far richer than the SERP suggests — but most of it sits in academic portals and biobanks rather than on Kaggle or Hugging Face. For sequencing work, Qiita, MGnify, GMrepo v3, and curatedMetagenomicData cover most use cases. For disease cohorts, IBDMDB and PRISM are the gold standard.

For Bristol-scale image ML, the public field is small and academic-on-request — and the underreported reality is that the largest annotated stool image collection in existence is operator-side, not academic. PoopCheck’s 180,000+ BSFS-annotated images from 25,000+ users is by itself larger than every public Bristol-labeled image dataset combined. Researchers planning ML work in this space should know that scale exists outside the open-data ecosystem, even if it is not currently distributed. The team that eventually publishes a permissively-licensed public BSFS dataset at scale will become the de-facto reference for this corner of the field.

Sources

  1. Qiita: rapid, web-enabled microbiome meta-analysis. Nature Methods, 2018.
  2. MGnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Research, 2023.
  3. GMrepo v2: a curated human gut microbiome database with special focus on disease markers and cross-dataset comparison. Nucleic Acids Research, 2022.
  4. An integrated catalog of reference genes in the human gut microbiome (Integrated Gene Catalog). Nature Biotechnology, 2014.
  5. Multi-omics of the gut microbial ecosystem in inflammatory bowel diseases (IBDMDB). Nature, 2019.
  6. The Integrative Human Microbiome Project. Nature, 2019.
  7. The Cancer Microbiome Atlas: a pan-cancer microbial reference resource. Cell Host & Microbe, 2020.
  8. 1000IBD: a multi-omics cohort study. BMC Gastroenterology, 2018.
  9. Human Stools Classification for Gastrointestinal Health using an Improved Deep Learning Architecture (Zhang et al.). CVPR Workshop, 2022.
  10. Smart Toilets for Stool Image Analysis (Zhou et al.). PMLR, 2021.
  11. Deep Learning Model Using Stool Pictures for Predicting Endoscopic Mucosal Inflammation in UC. American Journal of Gastroenterology, January 2025.
  12. Augmenting gastrointestinal health: A deep learning approach to human stool recognition (Hachuel et al.). Scientific Reports, 2021.
  13. SILVA in 2026: linking rRNA-based phylogeny with genome-derived taxonomy. Nucleic Acids Research, 2026.
  14. Greengenes2 unifies 16S and shotgun-based microbiome studies. Nature Biotechnology, 2023.
  15. The Earth Microbiome Project. UCSD / Knight Lab.
  16. The Dutch Microbiome Project on EGA.

Ready to track your gut health?

Download PoopCheck free and get your first AI stool analysis in seconds.

Scan your first photo free
P
PoopCheck Team

The PoopCheck team is dedicated to making digestive health tracking accessible, accurate, and private for everyone.

4.8 · Get the app free