CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01andlyu /Public-YAM-runs Public-YAM-runs Physical bimanual YAM episodes recorded by the BluPe operator station. Each run adds an episode to this repository. Failed, interrupted, stopped and timed-out runs are retained and labeled; these are not all successful demonstrations. A model saying done is not independently verified task success. Loading from datasets import load_dataset runs = load_dataset("andlyu/Public-YAM-runs", split="train") usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.image100K<n<1M2 likes12k downloads5h agoHugging Face02ruslanmv /sports-trends-dataset ⚽🏀🎾🏏 Sports-Trends Dataset A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits. The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day. TL;DR — A continuously-updated, medallion-architecture data lake for football, basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.tabulartabular-classificationn<1K5 likes4.5k downloads1h agoHugging Face03RussianNLP /coat Dataset Card for CoAT🧥 Dataset Description CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications. Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.tabulartext-classification100K<n<1M4 likes2.2k downloads10mo agoHugging Face04sewa-rural-care /anemia-survey-datasetgated Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India Dataset: sewa-rural-care/anemia-survey-dataset Contact: sewarural@ymail.com Version: 1.0 — July 2026 Dataset Summary This dataset supports research into non-invasive, smartphone-based anemia screening applicable to low-resource and rural healthcare settings. It was collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.tabularimage-classification1K<n<10K7 likes1.5k downloads3mo agoHugging Face05ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes1.3k downloads2y agoHugging Face06SaylorTwift /RULER-8192-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes1.3k downloads1y agoHugging Face07ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face08RUC-AIBOX /ICPC-Evaltabularn<1K3 likes952 downloads1y agoHugging Face09self-long /RULER-llama3-1M RULER-Llama3-1M A 1M token version of the RULER dataset based on the Llama-3 chat template. It is automatically generated based on the scripts available in the RULER repository: https://github.com/NVIDIA/RULER. It is designed for evaluating the performance of Long Language Models (LLMs) on various tasks with varying sequence lengths. How to Use from datasets import load_dataset LENGTH_IN_STRING = ['4k', '8k', '16k', '32k', '64k', '128k', '256k', '512k', '1M'] TASKS =… See the full description on the dataset page: https://huggingface.co/datasets/self-long/RULER-llama3-1M.tabular10K<n<100K3 likes898 downloads2y agoHugging Face10bicycleman15 /ruler-300-seed42 Frozen RULER 300, seed 42 This dataset freezes the exact RULER inputs used by the short-long-pretraining native evaluation suite. Repository: bicycleman15/ruler-300-seed42 Rows: 6,300 Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq Context lengths: 1024, 2048, 4096 Samples per task/length: 300 Seed: 42 Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.tabularquestion-answering1K<n<10K0 likes878 downloads1mo agoHugging Face11yjernite /prof_report__runwayml-stable-diffusion-v1-5__multi__24 Dataset Card for "prof_report__runwayml-stable-diffusion-v1-5__multi__24" More Information needed tabular1K<n<10K0 likes763 downloads3y agoHugging Face12rudymartin /georsct GeoRSCT A geospatial regression benchmark for evaluating representation–solver compatibility. GeoRSCT is a benchmark and evaluation framework for studying when geospatial model performance reflects solver quality versus target difficulty, spatial leakage, aggregation effects, scale sensitivity, or representation–solver mismatch. This release (version 24.0.1) includes 31,789 U.S. ZIP Code Tabulation Areas (ZCTAs), 106 columns spanning 33 ACS features, 37 geospatial enrichment… See the full description on the dataset page: https://huggingface.co/datasets/rudymartin/georsct.geospatialtabular-regression1M<n<10M0 likes737 downloads3d agoHugging Face13RussianNLP /rublimp RuBLiMP Dataset Description RuBLiMP, or Russian Benchmark of Linguistic Minimal Pairs, is the first diverse and large-scale benchmark of minimal pairs in Russian. RuBLiMP includes 45k minimal pairs of sentences that differ in grammaticality and isolate morphological, syntactic, or semantic phenomena. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/rublimp.tabular10K<n<100K4 likes727 downloads1y agoHugging Face14RUC-AIBOX /OlymMATH-eval OlymMATH Evaluation Results OlymMATH is a dataset we introduced in Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models by Haoxiang Sun, Yingqian Min, Zhipeng Chen, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, Lei Fang, and Ji-Rong Wen. You can find more information on GitHub and HuggingFace 🤗. We have made our evaluation results for the avg@{8, 64} and cons@{8, 64} metrics in this dataset publicly available for academic research… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/OlymMATH-eval.tabularquestion-answering100K<n<1M5 likes702 downloads1y agoHugging Face15ASSERT-KTH /RunBugRun-Final Original Dataset + Tokenized Data + (Buggy + Fixed Embedding Pairs) + Difference Embeddings Overview This repository contains 4 related datasets for training a transformation from buggy to fixed code embeddings: Datasets Included 1. Original Dataset (train-00000-of-00001.parquet) Description: Legacy RunBugRun Dataset Format: Parquet file with buggy-fixed code pairs, bug labels, and language Size: 456,749 samples Load with: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/RunBugRun-Final.tabular100K<n<1M0 likes677 downloads9mo agoHugging Face16jablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes540 downloads3mo agoHugging Face17SaylorTwift /RULER-32768-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes497 downloads1y agoHugging Face18iDRAMALab /iDRAMA-rumble-2024 Dataset Summary iDRAMA-rumble-2024 is a large-scale dataset of 6,735 podcast videos from Rumble, an alternative Youtube-like platform. Using state-of-the-art models, we extract information across three modalities: 1) text, 2) audio, and 3) video. We detail the methodology for extracting information from podcast videos in the paper and release a first-of-its-kind dataset including data from different modalities: Metadata: Details about podcast videos, e.g., channel name, video name… See the full description on the dataset page: https://huggingface.co/datasets/iDRAMALab/iDRAMA-rumble-2024.image100K<n<1M2 likes466 downloads2y agoHugging Face19TacVerse /xtac-umi-g1-insert-rubber-stopper Representative frames from TacVerse's bimanual demonstrations. Collected with XTac-UMI-G1 grippers, released as LeRobot datasets. This dataset was created using LeRobot. Explore this dataset with the LeRobot Dataset Viewer. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "xtac_umi_g1", "total_episodes": 10, "total_frames": 6204, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/xtac-umi-g1-insert-rubber-stopper.tabularrobotics1K<n<10K0 likes465 downloads13d agoHugging Face20RoboCOIN /AIRBOT_MMK2_storage_rubiks_cube_and_cupgated AIRBOT_MMK2_storage_rubiks_cube_and_cup 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_rubiks_cube_and_cup.tabularrobotics10K<n<100K0 likes464 downloads9mo agoHugging Face21makermods /first_test_run_20260720_125646This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/first_test_run_20260720_125646.tabularroboticsn<1K0 likes461 downloads2mo agoHugging Face22RudrakshNanavaty /earnings-call-data S&P 500 earnings episodes (2005–2025) Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/RudrakshNanavaty/earnings-call-data.tabulartext-classification100K<n<1M0 likes429 downloads6mo agoHugging Face23Hieuman /stihi_rutabular1M<n<10M0 likes399 downloads10mo agoHugging Face24RoboCOIN /R1_Lite_move_the_position_of_the_rubiks_cubegated R1_Lite_move_the_position_of_the_rubiks_cube 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: place pick grasp 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_rubiks_cube.tabularrobotics1K<n<10K0 likes393 downloads9mo agoHugging Face25andlyu /Public-MakerMods-SO101-runs bimanual_so101 run visualizer LeRobot v2.1 playback mirror for robot-3652c537a175cbae. Original images and telemetry remain in the shared archive. Failed runs are retained; these are not all successful demonstrations. tabular10K<n<100K0 likes386 downloads9d agoHugging Face26RoboCOIN /AIRBOT_MMK2_place_the_umbrella_and_the_rulergated AIRBOT_MMK2_place_the_umbrella_and_the_ruler 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_umbrella_and_the_ruler.tabularrobotics1K<n<10K0 likes377 downloads9mo agoHugging Face27JackHsieh /Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-1024.statml-arxivtabular1M<n<10M0 likes376 downloads5mo agoHugging Face28mlfoundations-dev /hero_run_4_math_codetabular1M<n<10M0 likes368 downloads1y agoHugging Face29rumbleFTW /prism-librispeech-train-100tabular10K<n<100K0 likes365 downloads1y agoHugging Face30zait-ai /RuHeritage-Corpus RuHeritage-Corpus 🇬🇧 English Description RuHeritage-Corpus is a high-quality, curated dataset of Russian classical literature, specifically designed for the pre-training and continued pre-training (CPT) of Large Language Models (LLMs). The corpus focuses on the Golden and Silver Ages of Russian literature, providing models with exposure to rich vocabulary, complex syntactic structures, and stylistically flawless Russian text, acting as a "quality anchor"… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RuHeritage-Corpus.tabulartext-generation1K<n<10K2 likes349 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.