CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K542 likes127k downloads2d agoHugging Face02uclanecl /NECL_GPUstabularn<1K10 likes17k downloads25m agoHugging Face03MrigLabIITRopar /GroMo25 GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting Dataset Summary GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.imageimage-classification100K<n<1M2 likes14k downloads6mo agoHugging Face04mrlbenchmarks /global-piqa-nonparallel Global PIQA Non-Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems. In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.imagequestion-answering10K<n<100K40 likes6.2k downloads4mo agoHugging Face05gneubig /aime-1983-2024 AIME Problem Set 1983-2024 Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) from 1983 to 2024. The AIME is a prestigious mathematics competition for high school students in the United States and Canada. Dataset Summary Source: Kaggle - AIME Problem Set 1983-2024 License: CC0: Public Domain Total Problems: 2,250 Years Covered: 1983 to 2024 Main Task: Mathematics Problem Solving… See the full description on the dataset page: https://huggingface.co/datasets/gneubig/aime-1983-2024.tabulartext-classificationn<1K21 likes6k downloads2y agoHugging Face06nmayorga7 /gpqa_diamondtabularn<1K0 likes5.3k downloads1y agoHugging Face07mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes3.8k downloads4mo agoHugging Face08AstraTeam /generated-csvstabular100K<n<1M0 likes2.7k downloads10mo agoHugging Face09Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face10gbionics /cmu-fbx CMU Motion Capture Library — FBX Format A large-scale library of 2,548 human motion capture animations in .fbx format, organized by subject/motion number. This dataset is derived from the CMU Graphics Lab Motion Capture Database and is intended for use in robotics, computer animation, game development, and machine learning research involving human motion. Source page: https://rancidmilk.itch.io/free-character-animations Dataset Summary This dataset contains over 2… See the full description on the dataset page: https://huggingface.co/datasets/gbionics/cmu-fbx.3d1K<n<10K10 likes2.4k downloads7mo agoHugging Face11gtak1 /panlex-meanings Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from https://panlex.org. Dataset Details Dataset Description This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis. Each language subset consists of expressions (words and phrases). Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/gtak1/panlex-meanings.tabulartranslation10M<n<100M0 likes2.3k downloads8mo agoHugging Face12MidiAndTheGang /simplified_grooveThis is a copy of the Magenta Groove dataset The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes. tabular1K<n<10K0 likes2.3k downloads2y agoHugging Face13mnemoraorg /usgs-global-earthquake-catalog USGS Global Earthquake Catalog Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service. Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including: Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event. Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.tabulartext-classification1M<n<10M1 likes2.2k downloads11mo agoHugging Face14google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K153 likes1.5k downloads4y agoHugging Face15hamzas /nba-games NBA Games Data This data is an updated version of the original NBA Games by Nathan Lauga. Data source Code Updated to: 2025-02-13 The dataset retains the original format and includes the following files: games.csv ‚Äì Summary of NBA games, including scores and team details. games_details.csv ‚Äì Detailed player statistics for each game. players.csv ‚Äì Player information. ranking.csv ‚Äì Daily NBA team rankings. teams.csv ‚Äì List of all NBA teams. tabular1M<n<10M2 likes1.5k downloads2y agoHugging Face16RaccoonOnion /gpqa-swaptabularquestion-answering1K<n<10K0 likes1.5k downloads1y agoHugging Face17gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.4k downloads9mo agoHugging Face18linxy97 /genhome3d-1280 GenHome3D-1280 1,280 validated household and spatial-design assets in USDZ format, organized across 64 categories. Explore the visual catalog · Browse the GitHub repository · Download the versioned release · Read the generation method Dataset summary Assets 1,280 Categories 64 Assets per category 20 Runtime format USDZ Units Meters Asset license CC BY 4.0 Technical validation 1,280/1,280 pass Package validation 1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.3d1K<n<10K1 likes1.4k downloads2mo agoHugging Face19AnimeshShaw /GenIaC-SecBench GenIaC-SecBench A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code (IaC) against a size-matched human baseline. Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code (arXiv:2608.28021) Code: https://github.com/AnimeshShaw/GenIaC-SecBench Why this dataset exists Prior evaluations of generated IaC report vulnerability counts for models only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.tabulartext-generation10K<n<100K1 likes1.3k downloads2d agoHugging Face20ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.3k downloads1y agoHugging Face21gelnesr /RelaxDB RelaxDB Dataset The RelaxDB and RelaxDB-CPMG datasets are curated data of relaxation-dispersion NMR data. This dataset was used to evaluate our model Dyna-1. Both the model and the datasets were introduced in our paper "Learning millisecond protein dynamics from what is missing in NMR spectra". This HF datasets hosts the files for the RelaxDB data. More information on analysis from the paper, evaluation of the dataset using Dyna-1, or the Dyna-1 model itself can be found on… See the full description on the dataset page: https://huggingface.co/datasets/gelnesr/RelaxDB.tabular1K<n<10K9 likes1.3k downloads27d agoHugging Face22openai /genebench-pro-public-package GeneBench-Pro Public Case Studies This repository contains public GeneBench-Pro case studies. It is the self-contained package intended for public distribution, including Hugging Face publication. Package Layout <repo-root>/ ├── .gitattributes ├── README.md ├── LICENSE ├── problems.csv ├── checksums.sha256 ├── manifest.json ├── reference_definitions.md ├── reference_grader.py └── problems/ └── <eval_id>/ ├── eval_config.json ├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.documentn<1K15 likes1.3k downloads3mo agoHugging Face23aai510-group1 /telco-customer-churn Dataset Card for Telco Customer Churn This dataset contains information about customers of a fictional telecommunications company, including demographic information, services subscribed to, location details, and churn behavior. This merged dataset combines the information from the original Telco Customer Churn dataset with additional details. Dataset Details Dataset Description This merged Telco Customer Churn dataset provides a comprehensive view of customer… See the full description on the dataset page: https://huggingface.co/datasets/aai510-group1/telco-customer-churn.tabulartabular-classification1K<n<10K15 likes1.3k downloads2y agoHugging Face24beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.3k downloads7mo agoHugging Face25sebastiandizon /genius-song-lyricstabular1M<n<10M38 likes1.1k downloads3y agoHugging Face26HydraLM /glaive_function_calling_v1_standardizedtabular100K<n<1M5 likes1.1k downloads3y agoHugging Face27wangyz1999 /GameplayQA GameplayQA: A Decision-Dense POV-Synced Multi-Video Understanding Benchmark of 3D Virtual Agents Yunzhe Wang   Runhui Xu   Kexin Zheng   Tianyi Zhang Jayavibhav N. Kogundi   Soham Hans   Volkan Ustun University of Southern California ACL 2026 Corresponding Author: yunzhewa@usc.edu Overview GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.tabularvideo-text-to-text1K<n<10K7 likes1k downloads4mo agoHugging Face28GD-ML /TransitLM TransitLM: Dataset Release & Evaluation Protocol Dataset Description TransitLM is a dataset for public transit route planning in Chinese urban environments, designed to support training and evaluation of language models that generate structured transit routes from origin-destination information. The full dataset covers four cities: Beijing, Shanghai, Shenzhen, and Chengdu, and includes coordinates, station sequences, transfer structure, line information, and route… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/TransitLM.tabulartext-generation100K<n<1M82 likes922 downloads4mo agoHugging Face29bowen-upenn /GeoGrid_Bench GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data? We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong spatial and temporal dependencies, and unique multimodal representations including tabular data, heatmaps, and geographic visualizations. To assess how foundation… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/GeoGrid_Bench.imagetable-question-answering10K<n<100K1 likes886 downloads1y agoHugging Face30GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes883 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.