datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Qwen2VL-Data
Introduction
This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.
Project page: https://victorwz.github.io/Open-Qwen2VL
Code: https://github.com/Victorwz/Open-Qwen2VL
Dataset
ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1
datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.opendataLanguage: English (current) · 中文
Representative frames from TacVerse's bimanual
demonstrations.
Collected with XTac-UMI-G1 grippers, released as LeRobot
datasets.
TacVerse Open Data
Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes,
370.2 hours, 40.0M frames, ~145 GB.
Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/).
Collection timestamps have been removed from titles and metadata.
Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.open-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.Open-MOPD-Data
Open-MOPD Data
This repository contains the training and evaluation data released with
Open-MOPD, including mixed-domain supervised fine-tuning data, the shared
RL/OPD prompt mixture, and six evaluation benchmarks.
Dataset contents
Configuration
Description
Examples
rl_prompt_mix
Shared math, code, and instruction-following prompts for RL and OPD
86,931
sft_openr1_math_93k
Math SFT data in a unified think-tag format
93,733
sft_ocr_50k
Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.italian-schools-opendatacrystallography-open-database
Crystallography Open Database (COD) — Full Snapshot
A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset.
Snapshot Details
Field
Value
Snapshot date
2026-07-06
Metadata fetched
2026-07-06 18:51 (UTC+2) — 533,486 entries
CIF files downloaded
2026-07-06 18:34–21:58 — 533,862 files
Total rows
533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.open_model_evolution_data
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy.
It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/mmpr/open_model_evolution_data.datasets-dependents
datasets metrics
This dataset contains metrics about the huggingface/datasets package.
Number of repositories in the dataset: 4997
Number of packages in the dataset: 215
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 22 packages that have more than 1000 stars.
There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.ODA-Math-460k
ODA-Math-460k
ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models.
🧠 Dataset Summary
Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.open-customs-data
Open Customs Data
Dataset Summary
This dataset contains 65,163,468 import transactions from the Bureau of Customs of the Philippines spanning from 2012 to 2025. It represents a comprehensive record of goods imported into the Philippines, including detailed information on product classifications, values, duties paid, countries of origin, and ports of entry.
The data was extracted from the Bureau of Customs' electronic2mobile (e2m) system and has been cleaned, standardized… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/open-customs-data.opendata
OpenData Consortium
Three open datasets exported from the OpenData Consortium data platform.
Config
Description
Primary format
companies
~103M global companies with firmographic attributes
Parquet
locations
~273M business locations with address and geo data
Parquet
people
~101M business contacts linked to companies
Parquet
Usage
from datasets import load_dataset
companies = load_dataset("OpenDataFoundation/opendata", "companies")
locations =… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataFoundation/opendata.MathLake
MathLake: A Large-Scale Mathematics Dataset
MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.yelp-open-dataset-businessyelp-open-dataset-reviewsSYNTHETIC-1-SFT-Data-Code_decontaminated
Dataset description
This dataset is the same as open-r1/SYNTHETIC-1-SFT-Data-Code decontaminated against the benchmark datasets.
The decontamination has been run using the script in huggingface/open-r1:
python scripts/decontaminate.py \
--dataset "open-r1/SYNTHETIC-1-SFT-Data-Code" \
-c
...
Removed 5 samples from 'aime_2025'
Removed 50 samples from 'math_500'
Removed 13234 samples from 'lcb'
Initial size: 62953, Final size: 49664
Chinese-Open-Datajapan-municipal-open-data-atlas-2026
Japan Municipal Open Data Atlas 2026
Release status: public release 2026.08.8. This release was approved
after the recorded provenance, reproducibility, and cross-platform checks passed.
Official statistics for every Japanese municipality, already joined, typed, and
documented — plus the name-matching table you would otherwise have to build
yourself before any of it is usable.
Also distributed as a Kaggle dataset mirror
with runnable notebooks, browsable without installation… See the full description on the dataset page: https://huggingface.co/datasets/yhay81/japan-municipal-open-data-atlas-2026.open-customs-data
Open Customs Data
Dataset Summary
This dataset contains 65,163,468 import transactions from the Bureau of Customs of the Philippines spanning from 2012 to 2025. It represents a comprehensive record of goods imported into the Philippines, including detailed information on product classifications, values, duties paid, countries of origin, and ports of entry.
The data was extracted from the Bureau of Customs' electronic2mobile (e2m) system and has been cleaned, standardized… See the full description on the dataset page: https://huggingface.co/datasets/Pilin2005/open-customs-data.mbgfnet-open-shell-3d-gw-dataset
Open-shell 3d transition-metal complexes: PBE0 + unrestricted G0W0 quasiparticle dataset
1,240 open-shell mononuclear 3d transition-metal complexes (Ti, V, Cr, Mn, Fe, Co, Ni; spin
multiplicity 1-6, including broken-symmetry open-shell singlets — see note below), each with a
spin-unrestricted DFT (UKS-PBE0/cc-pVDZ) and one-shot unrestricted G0W0@PBE0/cc-pVDZ (UGWAC)
quasiparticle-energy calculation, built to extend MBGF-Net
(Venturella, Li, Hillenbrand, Zhu, arXiv:2407.20384) —… See the full description on the dataset page: https://huggingface.co/datasets/primateria/mbgfnet-open-shell-3d-gw-dataset.open-grant-data
Open Grant Data
756,453 U.S. funders · 95,735 grant opportunities · a 7,525,377-edge who-funds-whom
grant graph · 1,810,300 grantee organizations — public domain (CC0), embeddings included.
The open alternative to paywalled funder databases ($1,500–2,000/yr subscriptions). Built
entirely from public sources — IRS 990 / 990-PF filings, grants.gov, and state & foundation
portals — and released with modern embeddings so it's AI-ready out of the box.
Snapshot: June 2026… See the full description on the dataset page: https://huggingface.co/datasets/qwntl-labs/open-grant-data.open-smtp-error-dataset
Open SMTP Error Dataset
Open SMTP Error Dataset is an English-language, machine-readable reference
package for SMTP enhanced-status knowledge and cautious operational
classification. Version 1.1.0 contains 92 stable knowledge records and a
separate, auditable catalog of 120 classification rules.
The package is a reference artifact, not a live provider-policy feed. It helps
with observability, support, parser testing, and bounded delivery operations;
it does not establish… See the full description on the dataset page: https://huggingface.co/datasets/blazalek/open-smtp-error-dataset.open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.Open_Reaction_Data
ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset
This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems.
The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks.
Dataset Structure
The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases.
It includes:
orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.open-synth-training-databixi_opendata
Bixi's OpenData Modelisation
Here is a huggingface dataset refined with AntoineGiraud/dbt_bixi_opendata dbt-core project
that loads & transform bixi OpenData thanks to DuckDB 🦆🚀
Viz' exploration ideas
I used Power BI to explore the transformed data offloaded to .parquet (~ 4.7 times lighter than .csv)
After the pandemic, Montrealers realy went back to bixi 🥳
Data sources
Bixi Rentals OpenData (link)
🚲 Rentals V1 : from 2014 to 2021
for… See the full description on the dataset page: https://huggingface.co/datasets/AntoineGiraud/bixi_opendata.statsbomb-open-data-shots
StatsBomb Open Data — Football Shots (xG)
88,023 shot-level football records extracted from StatsBomb Open Data,
spanning 67 years of football (1958–2025) across 21 competitions, 48 seasons, 308 teams, and 6,147 players.
Includes StatsBomb's own xG values as labels, making this the most complete open football shot dataset
available on HuggingFace for training Expected Goals models.
Highlights
🏆 Lionel Messi — 2,670 shots, the most of any player in the dataset (18 La… See the full description on the dataset page: https://huggingface.co/datasets/ClementeH/statsbomb-open-data-shots.statsbomb-open-data-shots
StatsBomb Open Data — Football Shots (xG)
88,023 shot-level football records extracted from StatsBomb Open Data,
spanning 67 years of football (1958–2025) across 21 competitions, 48 seasons, 308 teams, and 6,147 players.
Includes StatsBomb's own xG values as labels, making this the most complete open football shot dataset
available on HuggingFace for training Expected Goals models.
Highlights
🏆 Lionel Messi — 2,670 shots, the most of any player in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/FelixCS3/statsbomb-open-data-shots.yelp-open-dataset-businessyelp-open-dataset-top-businessesyelp-open-dataset-top-reviews-per-business
