datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.jpl-small-body-database
JPL Small-Body Database
Credit: NASA/ESA
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of all known asteroids and comets with orbital elements, physical parameters, and discovery metadata. Updated daily from NASA JPL.
The JPL Small-Body Database (SBDB) is the authoritative source for orbital and physical data on all known asteroids, comets, and other small bodies. It is maintained by the Solar System Dynamics group at… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/jpl-small-body-database.crystallography-open-database
Crystallography Open Database (COD) — Full Snapshot
A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset.
Snapshot Details
Field
Value
Snapshot date
2026-07-06
Metadata fetched
2026-07-06 18:51 (UTC+2) — 533,486 entries
CIF files downloaded
2026-07-06 18:34–21:58 — 533,862 files
Total rows
533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.code-databaseTrialPanorama-database
Quick start
The easiest way to download the dataset to your local is to use huggingface-cli. The specific command you can use is
huggingface-cli download zifeng-ai/TrialPanorama-database --local-dir LOCAL_DIR --repo-type dataset
where LOCAL_DIR should be replaced with the target directory you want to save your dataset to.
Update history
Aug.4 2025: updated tables with the full set of studies
Dataset website: https://ryanwangzf.github.io/projects/trialpanorama… See the full description on the dataset page: https://huggingface.co/datasets/TrialPanorama/TrialPanorama-database.ucs-satellite-database
UCS Satellite Database
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
The Union of Concerned Scientists (UCS) Satellite Database is the most comprehensive publicly available database of operational satellites. Updated roughly quarterly, it includes detailed information about each operational satellite: its name, country of registry, operator, purpose, orbital parameters, launch details, and physical characteristics.
What… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/ucs-satellite-database.fragrance-database
FragDB v5.16 — Fragrance Database (Multilingual Sample)
The most comprehensive structured fragrance database available. This is a free sample of FragDB: 140,230 perfumes, 23 languages — 10-row CSV samples at root.
Full dataset: fragdb.net.
What's New in v5.16
Data updated from v5.15 → v5.16 (snapshot 2026-09-19):
Fragrances: 139,501 → 140,230 (+729)
Brands: 8,272 → 8,316 (+44)
Perfumers: 3,116 → 3,126 (+10)
Notes: 2,596 → 2,606 rows in notes.csv (+10)
Companion… See the full description on the dataset page: https://huggingface.co/datasets/FragDBnet/fragrance-database.USDA-Phytochemical-Database-JSON
Ethno-API v2.4.0 — Public Sample
Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.free-global-stock-ticker-database
Free Global Stock Ticker Database
Global stocks and ETFs with listings, identifiers, aliases, and reviewed symbol changes. Maintained by Adanos Software GmbH for ticker detection, identifier resolution, and market-data workflows.
Source and project page: https://adanos.org
GitHub repository: https://github.com/adanos-software/free-ticker-database
Dataset package: adanosorg/free-global-stock-ticker-database
Version: 3.15.0
Contents
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adanosorg/free-global-stock-ticker-database.vpic-database
NHTSA vPIC Curated Vehicle Database
Curated vehicle database derived from NHTSA's vPIC
data, optimized for VIN decoding via the vin-decode-mcp server.
Source
Data sourced from the National Highway Traffic Safety Administration (NHTSA),
an agency of the United States Department of Transportation. NHTSA is a
government agency and the services provided are free for use by the public
as part of their Open Data initiative. No API key or registration required.
Download… See the full description on the dataset page: https://huggingface.co/datasets/joakes90/vpic-database.text-to-art-database
Vieutopia T2A Privacy Train v1
Dataset Summary
Privacy-safe text-to-image dataset repacked into Parquet shards with embedded image bytes.
Scope: text-to-image outputs only
Excluded: image-to-image pipelines (pix2pix_*, pst_*)
Privacy: no raw task UUIDs, no user/device fields
Storage format: parquet shards (image as binary bytes), no image_path dependency
Splits
samples
train: 117572
validation: 6532
test: 6532
total: 130636
iterations… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/text-to-art-database.slot-database
Slot Machine Database — 5,669 Slots from 58 Providers
Comprehensive dataset of online slot machine metadata covering 58 game providers. Each record includes RTP, volatility, max win multiplier, grid layout, mechanics, themes, features, and bet ranges.
Homepage: slot.report
API: slot.report/api/
Dataset Description
This dataset contains structured metadata for 5,669 online slot machines, making it one of the largest publicly available slot game databases.… See the full description on the dataset page: https://huggingface.co/datasets/slreport/slot-database.astronaut-database
Astronaut Database
Credit: NASA/GSFC/Suomi NPP
Part of a dataset collection on Hugging Face.
Dataset description
Complete database of every person who has traveled to space, sourced from Wikidata.
Since Yuri Gagarin's flight aboard Vostok 1 in April 1961, fewer than 700 individuals have crossed the Karman line (100 km altitude). This dataset records every one of them, from the Mercury Seven and Voskhod cosmonauts through Space Shuttle crews, ISS… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/astronaut-database.gpu-database
GPU Database
Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs.
2,824 GPUs across NVIDIA, AMD, and Intel
Part of RightNow — AI-powered code editor for GPU kernel development
Data
Vendor
GPUs
File
NVIDIA
1,286
data/nvidia/all.json
AMD
1,292
data/amd/all.json
Intel
180
data/intel/all.json
All
2,824
data/all-gpus.json
Schema
Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.database-agent-runs
LibreDB Agent Benchmark
8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls
This is the complete measurement record behind the paper What Stops a Small Language Model From
Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the
runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.MRC-psycholinguistic-database
MRC Psycholinguistic Database
This is the complete MRC psycholinguistic database as found on https://websites.psychology.uwa.edu.au/school/mrcdatabase/uwa_mrc.htm.
Usage
This dataset is ideal for training and evaluating machine learning models for English word concreteness.
Acknowledgments
We extend our heartfelt gratitude to all the authors of the original dataset.
License
This dataset is made available under the MIT license.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/nithishvenkat4/product-database.pima-indians-diabetes-database
Pima Indians Diabetes Dataset Split
This directory contains split datasets of Pima Indians Diabetes Database.
For each splits, we have
Mock data: The mock data is a smaller dataset (10 rows for both train and test) that is used to test the model and data processing code.
Private data: Each private data contains 123-125 rows for training, and 32-33 rows for testing.
openfoodfacts-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/HashimAliii/openfoodfacts-database.horse-racing-database-fr
Huggy Database — French Horse Races (1996-2026)
Official site: https://huggydatabase.com/ · Interactive Query Builder: https://huggydatabase.com/#query
356,371 races · 30 years (1996 → 2026) · 40+ variables per runner · SQL-ready.
A structured French horse-racing archive covering harness trot, mounted trot, flat and jumps, with PMU/PMH odds spread, shoeing, blinkers, going, autostart and pedigree.
🇬🇧 English
What is in this repository
This… See the full description on the dataset page: https://huggingface.co/datasets/huggy-engine/horse-racing-database-fr.space-agency-database
Space Agency Database
Credit: NASA/GSFC/Suomi NPP
Part of a dataset collection on Hugging Face.
Dataset description
Database of space agencies and related governmental space organizations worldwide, sourced from Wikidata.
From NASA and Roscosmos to emerging national programs in Asia, Africa, and Latin America, this dataset catalogs every governmental space agency and related intergovernmental organization known to Wikidata. It covers founding dates… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-agency-database.pima-indians-diabetes-database-partitions
Pima Indians Diabetes Dataset Split
This directory contains a dataset split for Pima Indians Diabetes Database.
Mock Data
The mock data is a smaller dataset (10 rows) that is used to test the model components.
Private Data
The private data is the remaining data that is used to train the model.
database-query-logs-synthetic
Database Query Logs (synthetic)
3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL
Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text,
type, complexity, execution timing, and row-count metadata.
These queries are synthetic
The queries were programmatically generated, not captured from production systems.
They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.spacecraft-database
Spacecraft Database
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Comprehensive database of spacecraft sourced from Wikidata — satellites, probes, space stations, and more — spanning the entire history of the Space Age.
From the earliest Sputnik satellites to modern mega-constellations and deep space probes, this dataset catalogs spacecraft across seven decades of spaceflight. Each record includes launch and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/spacecraft-database.openhands-commit-noise-databases
OpenHands Commit Noise Databases
This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%.
Each archive expands to noise_NNN/<repository>/ directories containing:
commits.db: SQLite commit records
commits.faiss: normalized inner-product FAISS index
commits.meta.jsonl: FAISS row-to-commit metadata
commits.index_meta.json: embedding and index configuration
The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.data_base_nerproduct-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/AmieTorin/product-database.Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.
