datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.jpl-small-body-database
JPL Small-Body Database
Credit: NASA/ESA
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of all known asteroids and comets with orbital elements, physical parameters, and discovery metadata. Updated daily from NASA JPL.
The JPL Small-Body Database (SBDB) is the authoritative source for orbital and physical data on all known asteroids, comets, and other small bodies. It is maintained by the Solar System Dynamics group at… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/jpl-small-body-database.crystallography-open-database
Crystallography Open Database (COD) — Full Snapshot
A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset.
Snapshot Details
Field
Value
Snapshot date
2026-07-06
Metadata fetched
2026-07-06 18:51 (UTC+2) — 533,486 entries
CIF files downloaded
2026-07-06 18:34–21:58 — 533,862 files
Total rows
533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.llmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.CTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.code-databaseTrialPanorama-database
Quick start
The easiest way to download the dataset to your local is to use huggingface-cli. The specific command you can use is
huggingface-cli download zifeng-ai/TrialPanorama-database --local-dir LOCAL_DIR --repo-type dataset
where LOCAL_DIR should be replaced with the target directory you want to save your dataset to.
Update history
Aug.4 2025: updated tables with the full set of studies
Dataset website: https://ryanwangzf.github.io/projects/trialpanorama… See the full description on the dataset page: https://huggingface.co/datasets/TrialPanorama/TrialPanorama-database.roadmap-databasesTORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/abnerh/TORGO-database.ucs-satellite-database
UCS Satellite Database
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
The Union of Concerned Scientists (UCS) Satellite Database is the most comprehensive publicly available database of operational satellites. Updated roughly quarterly, it includes detailed information about each operational satellite: its name, country of registry, operator, purpose, orbital parameters, launch details, and physical characteristics.
What… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/ucs-satellite-database.Telegram-Databasefragrance-database
FragDB v5.16 — Fragrance Database (Multilingual Sample)
The most comprehensive structured fragrance database available. This is a free sample of FragDB: 140,230 perfumes, 23 languages — 10-row CSV samples at root.
Full dataset: fragdb.net.
What's New in v5.16
Data updated from v5.15 → v5.16 (snapshot 2026-09-19):
Fragrances: 139,501 → 140,230 (+729)
Brands: 8,272 → 8,316 (+44)
Perfumers: 3,116 → 3,126 (+10)
Notes: 2,596 → 2,606 rows in notes.csv (+10)
Companion… See the full description on the dataset page: https://huggingface.co/datasets/FragDBnet/fragrance-database.GZ_IsoTech
Dataset Card for GZ_IsoTech Dataset
Original Content
The dataset is created and used for Guzheng playing technique detection by [1]. The original dataset comprises 2,824 variable-length audio clips showcasing various Guzheng playing techniques. Specifically, 2,328 clips were sourced from virtual sound banks, while 496 clips were performed by a professional Guzheng artist.
The clips are annotated in eight categories, with a Chinese pinyin and Chinese characters written in… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/GZ_IsoTech.erhu_playing_tech
Dataset Card for Erhu Playing Technique
Original Content
This dataset was created and has been utilized for Erhu playing technique detection by [1], which has not undergone peer review. The original dataset comprises 1,253 Erhu audio clips, all performed by professional Erhu players. These clips were annotated according to three levels, resulting in annotations for four, seven, and 11 categories. Part of the audio data is sourced from the CTIS dataset described earlier.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/erhu_playing_tech.USDA-Phytochemical-Database-JSON
Ethno-API v2.4.0 — Public Sample
Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.free-global-stock-ticker-database
Free Global Stock Ticker Database
Global stocks and ETFs with listings, identifiers, aliases, and reviewed symbol changes. Maintained by Adanos Software GmbH for ticker detection, identifier resolution, and market-data workflows.
Source and project page: https://adanos.org
GitHub repository: https://github.com/adanos-software/free-ticker-database
Dataset package: adanosorg/free-global-stock-ticker-database
Version: 3.15.0
Contents
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adanosorg/free-global-stock-ticker-database.vpic-database
NHTSA vPIC Curated Vehicle Database
Curated vehicle database derived from NHTSA's vPIC
data, optimized for VIN decoding via the vin-decode-mcp server.
Source
Data sourced from the National Highway Traffic Safety Administration (NHTSA),
an agency of the United States Department of Transportation. NHTSA is a
government agency and the services provided are free for use by the public
as part of their Open Data initiative. No API key or registration required.
Download… See the full description on the dataset page: https://huggingface.co/datasets/joakes90/vpic-database.Databasegdb_databasesfungi_trait_circus_database
fungi_trait_circus_database
大菌輪「Trait Circus」データセット(統制形質)
最終更新日:2025/09/28
重要:データ形式を大幅に更新しました(v2.0)
Languages
Japanese and English
Please do not use this dataset for academic purposes for the time being. (casual use only)
非専門家が作成したデータセットです。学術目的での使用はご遠慮ください。
更新履歴
2025/09/28 (v2.0) - データ構造を全面改訂、Parquet形式に移行、データ量を約2倍に拡充(約400万件)
2025/08/12 (v1.0) - 初回公開版(約180万件)
概要
Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_trait_circus_database.text-to-art-database
Vieutopia T2A Privacy Train v1
Dataset Summary
Privacy-safe text-to-image dataset repacked into Parquet shards with embedded image bytes.
Scope: text-to-image outputs only
Excluded: image-to-image pipelines (pix2pix_*, pst_*)
Privacy: no raw task UUIDs, no user/device fields
Storage format: parquet shards (image as binary bytes), no image_path dependency
Splits
samples
train: 117572
validation: 6532
test: 6532
total: 130636
iterations… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/text-to-art-database.osint-tool-databaseslot-database
Slot Machine Database — 5,669 Slots from 58 Providers
Comprehensive dataset of online slot machine metadata covering 58 game providers. Each record includes RTP, volatility, max win multiplier, grid layout, mechanics, themes, features, and bet ranges.
Homepage: slot.report
API: slot.report/api/
Dataset Description
This dataset contains structured metadata for 5,669 online slot machines, making it one of the largest publicly available slot game databases.… See the full description on the dataset page: https://huggingface.co/datasets/slreport/slot-database.In-stock-Database
Molport In-Stock Database
The Molport In-Stock Database contains SMILES strings and Molport IDs for all 6.1 million in-stock molecules, covering both screening compounds and building blocks that are currently available for purchase.
Contents
SMILES – The molecular structure in SMILES notation.
SMILES_CANONICAL – Canonicalized SMILES representation.
MOLPORTID – Unique Molport identifier for each compound.
Scope
This dataset includes:
All screening… See the full description on the dataset page: https://huggingface.co/datasets/molport/In-stock-Database.ACCOUNTING_DATABASESastronaut-database
Astronaut Database
Credit: NASA/GSFC/Suomi NPP
Part of a dataset collection on Hugging Face.
Dataset description
Complete database of every person who has traveled to space, sourced from Wikidata.
Since Yuri Gagarin's flight aboard Vostok 1 in April 1961, fewer than 700 individuals have crossed the Karman line (100 km altitude). This dataset records every one of them, from the Mercury Seven and Voskhod cosmonauts through Space Shuttle crews, ISS… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/astronaut-database.isbndb-full-database
Dataset Card for "isbndb-annas"
More Information needed
gpu-database
GPU Database
Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs.
2,824 GPUs across NVIDIA, AMD, and Intel
Part of RightNow — AI-powered code editor for GPU kernel development
Data
Vendor
GPUs
File
NVIDIA
1,286
data/nvidia/all.json
AMD
1,292
data/amd/all.json
Intel
180
data/intel/all.json
All
2,824
data/all-gpus.json
Schema
Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.
