Skip to content
PoopCheck PoopCheck
FOR BUSINESS

Annotated stool image dataset

185,000+ real-world stool images, multi-label annotated for Bristol Stool Scale class, color, and pathology. The same dataset our production models are trained on — available for your ML team or research group.

What's included

One of the largest consumer stool image datasets in existence, delivered with labels ready for training.

185,000+ labeled images

Real-world stool photos collected through the PoopCheck app, captured across devices, lighting, and demographics.

Bristol Stool Scale class

Every image labeled Type 1–7 on the medically recognized Bristol Stool Scale.

Color & pathology flags

Multi-label annotations: color category, mucus, blood indicators, consistency, and shape anomalies.

JSON labels + images

Standard delivery format, with custom schemas available. Paired with per-image metadata when applicable.

What teams do with it

Training CV models

Large, diverse, label-ready image set — skip years of collection and annotation.

Academic research

Research-license terms for universities and published studies on gut health and digestive medicine.

Clinical validation

Benchmark your pipeline against a standardized external dataset before deploying to patients.

Benchmark data

A real-world evaluation set for gut-health AI, not another lab-sanitized collection.

Annotation methodology

Labels are produced by trained annotators against the Bristol Stool Scale and a standardized color/pathology rubric, with a portion of the set double-labeled for QA. Disagreements are resolved by senior review. The same labeling pipeline feeds our production model's training data.

Licensing

Research licenses are available for academic use with publication rights. Commercial licenses cover product development, with terms scaled by subset size and use case. We'll send a draft agreement after an intro call.

Frequently asked questions

How is the dataset licensed?
We offer research licenses for academic use (with publication rights) and commercial licenses for product development. Terms vary by use case — contact us and we'll scope a license that fits.
How is the data delivered?
Default delivery is a signed S3 bucket with images and JSON label files, plus a schema document. We can accommodate custom formats (COCO, CSV, per-label folders) on request.
Can we get a custom subset?
Yes. Common requests include "only Types 3–5", "only flagged samples", balanced gender/age splits, or pediatric-only subsets. We can scope these from the master dataset.
How is the data de-identified?
Images contain no PHI by design — we collect only the stool photo and app-side analysis metadata. Our annotation and delivery pipelines include additional scrubbing to ensure no identifying information leaks through metadata or file names.
What are the usage restrictions?
Standard restrictions: no redistribution outside your organization, no resale, and no use that would enable re-identification. Full terms are in the license agreement we send before delivery.
How is it priced?
Pricing depends on subset size, annotation depth, and whether the license is research or commercial. Contact us with the specifics of your need and we'll send a quote.

Also available

Request dataset access

Tell us about your use case and any specific subset needs. We'll reply within one business day.

0 / 2000

Or email us directly at marco@softallthings.com.

4.8 · Get the app free