ThorKl/theobroma
THEOBROMA v1.36 An aggregated open database of 1,132,805 natural products from 29 sources, with per-compound license auditing, three-tier classification provenance, and stereochemistry-aware deduplication. Live instance: https://theobroma.l3s.uni-hannover.de Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest) This release: https://doi.org/10.5281/zenodo.22816330 Preprint: https://doi.org/10.64898/2026.06.12.731585 Licensing… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/theobroma.
THEOBROMA v1.36
An aggregated open database of 1,132,805 natural products from 29 sources, with per-compound license auditing, three-tier classification provenance, and stereochemistry-aware deduplication.
- Live instance: https://theobroma.l3s.uni-hannover.de
- Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest)
- This release: https://doi.org/10.5281/zenodo.22816330
- Preprint: https://doi.org/10.64898/2026.06.12.731585
Licensing
This dataset is not uniformly licensed. Each compound carries a resolved license tier in the license_tier column, computed under a most-restrictive-wins rule across all attesting sources.
896,242 compounds (79.12%) are open-licensed under the most restrictive resolution. A least-restrictive bound is also recorded in tier_rank_min, under which 1,029,587 compounds (90.9%) reach CC BY or better from at least one attesting source. Downstream use must be governed by license_tier per compound, not by any single license applied to the whole dataset. The license_attestations table records which source asserted which terms for each compound, and source_licenses maps each source to its tier.
Tables
Note that license_tier appears in two tables with different meanings: in compounds it is the resolved tier after most-restrictive-wins across all attesting sources, and in license_attestations it is the tier stated by that individual source. Use suffixes when joining the two.
Nested columns are stored as Arrow structs rather than JSON strings: chebi_xrefs, the three lineage columns in compound_taxonomy, and native_distribution. secondary_kingdoms in resolved_taxonomy is a list of strings.
