CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01weizhiwang /Open-Qwen2VL-Data Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources. Project page: https://victorwz.github.io/Open-Qwen2VL Code: https://github.com/Victorwz/Open-Qwen2VL Dataset ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1 datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.imageimage-text-to-text10M<n<100M25 likes12k downloads1y agoHugging Face02TacVerse /opendataLanguage: English (current) · 中文 Representative frames from TacVerse's bimanual demonstrations. Collected with XTac-UMI-G1 grippers, released as LeRobot datasets. TacVerse Open Data Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes, 370.2 hours, 40.0M frames, ~145 GB. Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/). Collection timestamps have been removed from titles and metadata. Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.tabularrobotics10M<n<100M7 likes9.5k downloads7d agoHugging Face03ShawnChamberlain /open-economic-quant-research-data Open Economic & Quant Research Data Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation. Repository structure CasualLab/: causal inference and policy-simulation research content. Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.documenttabular-classificationn<1K0 likes2.2k downloads1mo agoHugging Face04BytedTsinghua-SIA /Open-MOPD-Data Open-MOPD Data This repository contains the training and evaluation data released with Open-MOPD, including mixed-domain supervised fine-tuning data, the shared RL/OPD prompt mixture, and six evaluation benchmarks. Dataset contents Configuration Description Examples rl_prompt_mix Shared math, code, and instruction-following prompts for RL and OPD 86,931 sft_openr1_math_93k Math SFT data in a unified think-tag format 93,733 sft_ocr_50k Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.tabulartext-generation1M<n<10M2 likes1.7k downloads1mo agoHugging Face05diatribe00 /italian-schools-opendatatabular10M<n<100M1 likes1.4k downloads2mo agoHugging Face06LMucko /crystallography-open-database Crystallography Open Database (COD) — Full Snapshot A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset. Snapshot Details Field Value Snapshot date 2026-07-06 Metadata fetched 2026-07-06 18:51 (UTC+2) — 533,486 entries CIF files downloaded 2026-07-06 18:34–21:58 — 533,862 files Total rows 533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.tabulartext-classification100K<n<1M0 likes1.3k downloads3mo agoHugging Face07mmpr /open_model_evolution_data Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy. It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/mmpr/open_model_evolution_data.tabular100K<n<1M6 likes1.1k downloads21h agoHugging Face08open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes554 downloads2y agoHugging Face09OpenDataArena /ODA-Math-460k ODA-Math-460k ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models. 🧠 Dataset Summary Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.tabularquestion-answering100K<n<1M105 likes463 downloads8mo agoHugging Face10bettergovph /open-customs-data Open Customs Data Dataset Summary This dataset contains 65,163,468 import transactions from the Bureau of Customs of the Philippines spanning from 2012 to 2025. It represents a comprehensive record of goods imported into the Philippines, including detailed information on product classifications, values, duties paid, countries of origin, and ports of entry. The data was extracted from the Bureau of Customs' electronic2mobile (e2m) system and has been cleaned, standardized… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/open-customs-data.tabular10M<n<100M2 likes426 downloads8mo agoHugging Face11OpenDataFoundation /opendata OpenData Consortium Three open datasets exported from the OpenData Consortium data platform. Config Description Primary format companies ~103M global companies with firmographic attributes Parquet locations ~273M business locations with address and geo data Parquet people ~101M business contacts linked to companies Parquet Usage from datasets import load_dataset companies = load_dataset("OpenDataFoundation/opendata", "companies") locations =… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataFoundation/opendata.tabular100M<n<1B1 likes378 downloads1mo agoHugging Face12OpenDataArena /MathLake MathLake: A Large-Scale Mathematics Dataset MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.tabularquestion-answering1M<n<10M21 likes351 downloads5mo agoHugging Face13yashraizad /yelp-open-dataset-businesstabular100K<n<1M1 likes350 downloads3y agoHugging Face14yashraizad /yelp-open-dataset-reviewstabular1M<n<10M0 likes288 downloads3y agoHugging Face15open-r1 /SYNTHETIC-1-SFT-Data-Code_decontaminated Dataset description This dataset is the same as open-r1/SYNTHETIC-1-SFT-Data-Code decontaminated against the benchmark datasets. The decontamination has been run using the script in huggingface/open-r1: python scripts/decontaminate.py \ --dataset "open-r1/SYNTHETIC-1-SFT-Data-Code" \ -c ... Removed 5 samples from 'aime_2025' Removed 50 samples from 'math_500' Removed 13234 samples from 'lcb' Initial size: 62953, Final size: 49664 tabular10K<n<100K3 likes288 downloads2y agoHugging Face16DistressedModel /Chinese-Open-Datatabular10M<n<100M0 likes245 downloads7mo agoHugging Face17yhay81 /japan-municipal-open-data-atlas-2026 Japan Municipal Open Data Atlas 2026 Release status: public release 2026.08.8. This release was approved after the recorded provenance, reproducibility, and cross-platform checks passed. Official statistics for every Japanese municipality, already joined, typed, and documented — plus the name-matching table you would otherwise have to build yourself before any of it is usable. Also distributed as a Kaggle dataset mirror with runnable notebooks, browsable without installation… See the full description on the dataset page: https://huggingface.co/datasets/yhay81/japan-municipal-open-data-atlas-2026.tabular100K<n<1M0 likes217 downloads2mo agoHugging Face18Pilin2005 /open-customs-data Open Customs Data Dataset Summary This dataset contains 65,163,468 import transactions from the Bureau of Customs of the Philippines spanning from 2012 to 2025. It represents a comprehensive record of goods imported into the Philippines, including detailed information on product classifications, values, duties paid, countries of origin, and ports of entry. The data was extracted from the Bureau of Customs' electronic2mobile (e2m) system and has been cleaned, standardized… See the full description on the dataset page: https://huggingface.co/datasets/Pilin2005/open-customs-data.tabular10M<n<100M0 likes200 downloads5mo agoHugging Face19primateria /mbgfnet-open-shell-3d-gw-dataset Open-shell 3d transition-metal complexes: PBE0 + unrestricted G0W0 quasiparticle dataset 1,240 open-shell mononuclear 3d transition-metal complexes (Ti, V, Cr, Mn, Fe, Co, Ni; spin multiplicity 1-6, including broken-symmetry open-shell singlets — see note below), each with a spin-unrestricted DFT (UKS-PBE0/cc-pVDZ) and one-shot unrestricted G0W0@PBE0/cc-pVDZ (UGWAC) quasiparticle-energy calculation, built to extend MBGF-Net (Venturella, Li, Hillenbrand, Zhu, arXiv:2407.20384) —… See the full description on the dataset page: https://huggingface.co/datasets/primateria/mbgfnet-open-shell-3d-gw-dataset.tabular1K<n<10K1 likes194 downloads9d agoHugging Face20qwntl-labs /open-grant-data Open Grant Data 756,453 U.S. funders · 95,735 grant opportunities · a 7,525,377-edge who-funds-whom grant graph · 1,810,300 grantee organizations — public domain (CC0), embeddings included. The open alternative to paywalled funder databases ($1,500–2,000/yr subscriptions). Built entirely from public sources — IRS 990 / 990-PF filings, grants.gov, and state & foundation portals — and released with modern embeddings so it's AI-ready out of the box. Snapshot: June 2026… See the full description on the dataset page: https://huggingface.co/datasets/qwntl-labs/open-grant-data.tabular10M<n<100M0 likes177 downloads3mo agoHugging Face21blazalek /open-smtp-error-dataset Open SMTP Error Dataset Open SMTP Error Dataset is an English-language, machine-readable reference package for SMTP enhanced-status knowledge and cautious operational classification. Version 1.1.0 contains 92 stable knowledge records and a separate, auditable catalog of 120 classification rules. The package is a reference artifact, not a live provider-policy feed. It helps with observability, support, parser testing, and bounded delivery operations; it does not establish… See the full description on the dataset page: https://huggingface.co/datasets/blazalek/open-smtp-error-dataset.textn<1K1 likes165 downloads2mo agoHugging Face22SDSC /open-pulse-hackathon-data-analysis LauzHack Projects Dataset Dataset Summary This dataset contains comprehensive information about projects submitted to LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project includes details about the project title, description, team members, awards, and categories. LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.tabularn<1K0 likes137 downloads5mo agoHugging Face23Azzindani /Open_Reaction_Data ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems. The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks. Dataset Structure The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases. It includes: orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.tabulartext-generation1M<n<10M0 likes130 downloads7mo agoHugging Face24tensorlink-dev /open-synth-training-datatabular10M<n<100M0 likes93 downloads7mo agoHugging Face25AntoineGiraud /bixi_opendata Bixi's OpenData Modelisation Here is a huggingface dataset refined with AntoineGiraud/dbt_bixi_opendata dbt-core project that loads & transform bixi OpenData thanks to DuckDB 🦆🚀 Viz' exploration ideas I used Power BI to explore the transformed data offloaded to .parquet (~ 4.7 times lighter than .csv) After the pandemic, Montrealers realy went back to bixi 🥳 Data sources Bixi Rentals OpenData (link) 🚲 Rentals V1 : from 2014 to 2021 for… See the full description on the dataset page: https://huggingface.co/datasets/AntoineGiraud/bixi_opendata.tabular1M<n<10M0 likes88 downloads10mo agoHugging Face26ClementeH /statsbomb-open-data-shots StatsBomb Open Data — Football Shots (xG) 88,023 shot-level football records extracted from StatsBomb Open Data, spanning 67 years of football (1958–2025) across 21 competitions, 48 seasons, 308 teams, and 6,147 players. Includes StatsBomb's own xG values as labels, making this the most complete open football shot dataset available on HuggingFace for training Expected Goals models. Highlights 🏆 Lionel Messi — 2,670 shots, the most of any player in the dataset (18 La… See the full description on the dataset page: https://huggingface.co/datasets/ClementeH/statsbomb-open-data-shots.tabulartabular-classification10K<n<100K0 likes82 downloads5mo agoHugging Face27FelixCS3 /statsbomb-open-data-shots StatsBomb Open Data — Football Shots (xG) 88,023 shot-level football records extracted from StatsBomb Open Data, spanning 67 years of football (1958–2025) across 21 competitions, 48 seasons, 308 teams, and 6,147 players. Includes StatsBomb's own xG values as labels, making this the most complete open football shot dataset available on HuggingFace for training Expected Goals models. Highlights 🏆 Lionel Messi — 2,670 shots, the most of any player in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/FelixCS3/statsbomb-open-data-shots.tabulartabular-classification10K<n<100K0 likes78 downloads3mo agoHugging Face28yashraizada /yelp-open-dataset-businesstabular100K<n<1M0 likes68 downloads3y agoHugging Face29yashraizad /yelp-open-dataset-top-businessestabular10K<n<100K0 likes67 downloads3y agoHugging Face30yashraizada /yelp-open-dataset-top-reviews-per-businesstabular10K<n<100K0 likes66 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.