CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-mll /glue Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.tabulartext-classification1M<n<10M1.1k likes843k downloads3y agoHugging Face02nvidia /HelpSteer2 HelpSteer2: Open-source dataset for training top-performing reward models HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses. This dataset has been created in partnership with Scale AI. When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.tabular10K<n<100K456 likes142k downloads2y agoHugging Face03neigezhu /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M12 likes64k downloads1mo agoHugging Face04ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes44k downloads9d agoHugging Face05wayslab /llm-network-study-data LLM-Network-Study-Data Per-request network captures (.pcapng) collected by the LLM-Network-Study benchmark harness (benchmark.py and the per-workload test scripts). Each directory holds one capture file per request, named request_<id>_run<n>_<timestamp>.pcapng. A directory name encodes four dimensions: <capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results Dimension legend Dimension Values Meaning Capture env ethernet Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.tabularn<1K0 likes32k downloads1mo agoHugging Face06nvidia /PhysicalAI-Robotics-GR00T-Teleop-Sim Simulation GR1 Tabletop Task 1K Dataset Dataset Description: The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands. This dataset is ready for non-commercial use. Dataset Owner(s): NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.tabular1M<n<10M19 likes19k downloads9mo agoHugging Face07nvidia /PhysicalAI-Robotics-GR00T-Teleop-GR1 Introduction TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments. Project page: https://dreamdojo-world.github.io/ Paper: https://arxiv.org/abs/2602.06949 Code: https://github.com/NVIDIA/DreamDojo How to Use Check out https://github.com/NVIDIA/DreamDojo Citation @article{gao2026dreamdojo, title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.tabular1M<n<10M28 likes18k downloads7mo agoHugging Face08NoeFlandre /osm-polygon-selection osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions like Brazilian states, Chinese provinces, Indian zones, US states, Canadian provinces, Japanese regions, and Indonesian islands — classified by size bin (small / medium / large, area in [0.1, 100] km²) and tagged by continent (Natural Earth admin0 lookup). Size bins: small — area in [0.1, 1) km² (10,000 m² to 1 km², roughly 100 m × 100 m… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-selection.tabularother10M<n<100M0 likes17k downloads3mo agoHugging Face09ITMO-NSS /Aiice Dataset Aiice benchmark dataset for Arctic sea ice concentration (SIC) forecasting, based on OSI-SAF satellite products (CC BY 4.0). Coverage Period: October 1978 – April 2026 Resolution: 25 km spatial, daily temporal Grid: 432×432 (Lambert Azimuthal Equal Area, EPSG:6931) Source products Product Source Period OSI-450-a SMMR, SSM/I, SSMIS 1978–2020 OSI-430-a SSMIS 2021–Jul 2025 OSI-438 AMSR2 Jul 2025–present… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/Aiice.tabularn<1K1 likes17k downloads5d agoHugging Face10MU-NLPC /Calc-asdiv_a Dataset Card for Calc-asdiv_a Summary The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo). The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily parsed (e.g. by BeautifulSoup). The data contains 3 types of tags: gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.tabular1K<n<10K1 likes16k downloads3y agoHugging Face11uclanecl /NECL_GPUstabularn<1K10 likes15k downloads20m agoHugging Face12nguha /legalbench Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.tabulartext-classification10K<n<100K188 likes15k downloads6mo agoHugging Face13cvlab /new-york-smells New York Smells: A Large Multimodal Dataset for Olfaction While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory data collected in natural settings. We present New York Smells, a large-scale dataset of paired image and olfactory signals captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.image10K<n<100K1 likes13k downloads2mo agoHugging Face14NJU-LINK /CodeTraceBenchCodeTraceBench A Benchmark for Agent Trajectory Diagnosis CodeTraceBench is a large-scale benchmark of 4,316 agent trajectories with human-verified step-level annotations for evaluating trajectory diagnosis systems. Each trajectory records the full action-observation sequence of a coding agent, annotated with incorrect and unuseful step labels. Part of the CodeTracer project — a self-evolving agent trajectory diagnosis system. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/CodeTraceBench.tabulartext-generation1K<n<10K7 likes13k downloads6mo agoHugging Face15nexar-ai /nexar_collision_predictiongated Nexar Collision Prediction Dataset This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle. Dataset The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.tabularvideo-classification1K<n<10K21 likes12k downloads3d agoHugging Face16NRVBench /nrvbench-review NR Video Editing Benchmark This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions. The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.imagevideo-to-videon<1K1 likes9.5k downloads5mo agoHugging Face17NoeFlandre /osm-polygon-wikidata-only OSM Polygon Wikidata, Wikipedia and Wikivoyage OSM polygons carrying wikidata=*, enriched with multilingual Wikipedia and Wikivoyage documents. The published tables preserve regional records and provenance. Source code: GitHub repository. Dataset snapshot Metric Value Polygon rows across regional extracts 1,184,110 Unique polygon identities (osm_type, osm_id) 1,157,841 Polygons with successful non-empty text (unique OSM identities) 650,663… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-only.image10M<n<100M1 likes9k downloads6d agoHugging Face18nebius /SWE-rebench-leaderboard Dataset Summary ❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks. These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.tabular1K<n<10K30 likes8.3k downloads2mo agoHugging Face19ambrosfitz /19c_newspapers_images_altotabular100K<n<1M4 likes8.2k downloads3mo agoHugging Face20nhblk123 /helaxai_data_pluse 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/nhblk123/helaxai_data_pluse.tabulartext-generation100M<n<1B1 likes8k downloads27d agoHugging Face21nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B129 likes7.3k downloads11mo agoHugging Face22rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes7.2k downloads7mo agoHugging Face23IPEC-COMMUNITY /libero_spatial_no_noops_1.0.0_lerobottabular10K<n<100K5 likes6.8k downloads11mo agoHugging Face24PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.5k downloads3y agoHugging Face25nebula /FakeCOCO FakeCOCO dataset Using 10 SOTA text-to-image models to generate fake images based on COCO captions over 1M images These models include: SD15 SD21 SDXL SD3 Playground2.5 PixArt alpha PixArt sigma unidiffuser Flux.1 Stable Cascade tabularimage-classification100K<n<1M0 likes6.2k downloads2y agoHugging Face26nlile /24-game Math Twenty Four (24s Game) Dataset A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012. In each puzzle, players must use exactly four numbers and basic arithmetic operations (+, -, ×, /) to… See the full description on the dataset page: https://huggingface.co/datasets/nlile/24-game.tabularmultiple-choice1K<n<10K14 likes6.2k downloads2y agoHugging Face27IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6.1k downloads11mo agoHugging Face28IPEC-COMMUNITY /libero_object_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6k downloads11mo agoHugging Face29IPEC-COMMUNITY /libero_10_no_noops_1.0.0_lerobottabular100K<n<1M3 likes5.9k downloads11mo agoHugging Face30liuhyuu /NetEaseCrowd 🧑‍🤝‍🧑 NetEaseCrowd: A Dataset for Long-term and Online Crowdsourcing Truth Inference View it in GitHub Introduction We introduce NetEaseCrowd, a large-scale crowdsourcing annotation dataset based on a mature Chinese data crowdsourcing platform of NetEase Inc.. NetEaseCrowd dataset contains about 2,400 workers, 1,000,000 tasks, and 6,000,000 annotations between them, where the annotations are collected in about 6 months. In this dataset, we provide ground truths for… See the full description on the dataset page: https://huggingface.co/datasets/liuhyuu/NetEaseCrowd.tabular1M<n<10M1 likes5.9k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.