datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.VSRQA_DatabaseFolder Description:
GT -- original Full HD uncompressed videos. Folder corresponding to each video contains all frames of this video
GTCompressed -- videos from "GT" folder with reduced resolution and compressed with different video codecs ("No SR Track"). Folder corresponding to each of the video contains all frames of this compressed video. Folders are named in the following format: "<sequence name>___<scale factor>_<codec name>_<values of target bitrate or qp>"
GTCompressedFullHD -- the… See the full description on the dataset page: https://huggingface.co/datasets/Divotion/VSRQA_Database.database-10k-val
database-10k-val
最终 40K 数据集的 10,000 样本验证子集(2026-09-17)。
选择约束
不含 inset:所有样本 layout_type != inset。
不含冻结/删除族:样本主 panel 与全部副 panel 的 chart_family 均不属于 area、matrix、set_relation。
从最终修复后的 40K 中按 (family, subtype, layout, difficulty) 分层、确定性抽取 10,000 条。
文件
figure2data_10k_val.sqlite3:子集 SQLite,10,000 samples / 70,000 documents。
images/:10,000 PNG。
documents/:10,000 JSON。
arrays/:有原始数组的样本 NPZ。
plans/generation_plan_10k_val.jsonl:10,000 行子集计划。… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/database-10k-val.Chess-AI-Database
chess-AI-database
This is the main database for Chess AI
For more detail, please see Chess-AI-Pytorch
figure2data-database-v2
figure2data databasev2 — 40K v6 合成科研图表数据集(最终交付版)
生成日期:2026-09-17 生成器:generator 1.6.0 / dataset_generation_revision v6(含两次 hotfix)
规模:40,000 样本(10 图族 / 33 亚型;area、matrix 冻结不生成,set_relation 已删除)
目录结构
databasev2/
├── figure2data.sqlite3 # 主数据库(2.0 GB:40,000 samples / 280,000 documents)
├── schema/ # sqlite schema
├── shards/ # 数据资产(shard = (样本序号-1)//1000)
│ └── shard_000 .. shard_039/
│ ├── images/ # PNG… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/figure2data-database-v2.lekuku-databaseaix-lichess-database
Aix-compatible Lichess database
This dataset contains the Lichess database of 7+ billion chess games, but in a format queryable using Aix, with the aixchess extension for DuckDB.
Three compression levels (Low, Medium, and High) are offered for the chess game encoding, enabling a trade-off between decoding speed and size. The Parquet files themselves are all zstd-compressed at level 19.
Generated using pgn-to-aix.
See my blog post for more details.
Example DuckDB commands:
INSTALL… See the full description on the dataset page: https://huggingface.co/datasets/thomasd1/aix-lichess-database.9router-databaseleaderboard_database
HAKARI-Bench Leaderboard Database
This dataset hosts the DuckDB database used by the
HAKARI-Bench leaderboard.
It is derived from the raw benchmark result artifacts in
hakari-bench/results
and is packaged for leaderboard, viewer, notebook, and SQL use.
The database is produced by the HAKARI-Bench implementation in
hakari-bench/hakari-bench.
Because the benchmark code, schema, and build workflow evolve over time, this
dataset card intentionally points to the canonical… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/leaderboard_database.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.trip-plus-databasef2d-database-v2
figure2data databasev2 — 40K v6 合成科研图表数据集(最终交付版)
生成日期:2026-09-17 生成器:generator 1.6.0 / dataset_generation_revision v6(含两次 hotfix)
规模:40,000 样本(10 图族 / 33 亚型;area、matrix 冻结不生成,set_relation 已删除)
目录结构
databasev2/
├── figure2data.sqlite3 # 主数据库(2.0 GB:40,000 samples / 280,000 documents)
├── schema/ # sqlite schema
├── shards/ # 数据资产(shard = (样本序号-1)//1000)
│ └── shard_000 .. shard_039/
│ ├── images/ # PNG… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/f2d-database-v2.database_exportjpl-small-body-database
JPL Small-Body Database
Credit: NASA/ESA
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of all known asteroids and comets with orbital elements, physical parameters, and discovery metadata. Updated daily from NASA JPL.
The JPL Small-Body Database (SBDB) is the authoritative source for orbital and physical data on all known asteroids, comets, and other small bodies. It is maintained by the Solar System Dynamics group at… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/jpl-small-body-database.crystallography-open-database
Crystallography Open Database (COD) — Full Snapshot
A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset.
Snapshot Details
Field
Value
Snapshot date
2026-07-06
Metadata fetched
2026-07-06 18:51 (UTC+2) — 533,486 entries
CIF files downloaded
2026-07-06 18:34–21:58 — 533,862 files
Total rows
533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.Iris_Database
Synthetic Iris Image Dataset
Overview
This repository contains a dataset of synthetic colored iris images generated using diffusion models based on our paper "Synthetic Iris Image Generation Using Diffusion Networks." The dataset comprises 17,695 high-quality synthetic iris images designed to be biometrically unique from the training data while maintaining realistic iris pigmentation distributions. In this repository we contain about 10000 filtered iris images with the… See the full description on the dataset page: https://huggingface.co/datasets/fatdove/Iris_Database.HITECH_DATABASEfukinoto-databaseh-rag_databasellmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.CTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.spider-databases
Spider Databases (SQLite)
A re-host of the SQLite databases from the Spider
text-to-SQL benchmark (Yu et al., 2018), packaged as spider_data.zip so it can be pinned by
commit SHA and verified by checksum.
The standard xlangai/spider parquet ships only the question/SQL pairs — not the databases
needed to execute queries. This repo fills that gap for reproducible execution-based evaluation.
Contents: spider_data/database/<db_id>/<db_id>.sqlite, spider_data/tables.json, and the… See the full description on the dataset page: https://huggingface.co/datasets/HAL-9001/spider-databases.steam-databasecy-database
cy-database — Calabi–Yau geometries for string-compactification workflows
Precomputed Calabi–Yau threefold data, organised as a family of sub-datasets, accessed through the stringforge infrastructure package. Each sub-dataset covers one class of Calabi–Yau constructions and is keyed by a class-specific identifier system.
This top-level card describes the conventions, layout, and loading interface that are common to all sub-datasets. Each sub-dataset has its own card with… See the full description on the dataset page: https://huggingface.co/datasets/aschachner/cy-database.dol-visas-database
DOL Visas Database (H-1B LCA + PERM)
Every H-1B/H-1B1/E-3 Labor Condition Application and PERM permanent labor
certification application disclosed by the DOL Office of Foreign Labor
Certification, FY2015 to present, as a single queryable DuckDB database.
8,812,639 rows across 2 tables.
Table
Description
Row Count
Column Count
Date Range
lca
H-1B/H-1B1/E-3 Labor Condition Applications, one row per application per disclosure file, FY2015-present
7,479,697
110
FY2015 to… See the full description on the dataset page: https://huggingface.co/datasets/Nason/dol-visas-database.HITECH_DATABASEE-Waste-Databasepianos
Dataset Card for Piano Sound Quality Dataset
The original dataset is sourced from the Piano Sound Quality Dataset, which includes 12 full-range audio files in .wav/.mp3/.m4a format representing seven models of pianos: Kawai upright piano, Kawai grand piano, Young Change upright piano, Hsinghai upright piano, Grand Theatre Steinway piano, Steinway grand piano, and Pearl River upright piano. Additionally, there are 1,320 split monophonic audio files in .wav/.mp3/.m4a format, bringing… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/pianos.code-databaseSkyWorld-OSM-Database
