datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Public-YAM-runs
Public-YAM-runs
Physical bimanual YAM episodes recorded by the BluPe operator station.
Each run adds an episode to this repository. Failed, interrupted, stopped and
timed-out runs are retained and labeled; these are not all successful demonstrations.
A model saying done is not independently verified task success.
Loading
from datasets import load_dataset
runs = load_dataset("andlyu/Public-YAM-runs", split="train")
usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.sports-trends-dataset
⚽🏀🎾🏏 Sports-Trends Dataset
A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits.
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day.
TL;DR — A continuously-updated, medallion-architecture data lake for football,
basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an
engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.coat
Dataset Card for CoAT🧥
Dataset Description
CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications.
Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.RULER-8192-Qwen2.5-3B-tokenizerru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.ICPC-EvalRULER-llama3-1M
RULER-Llama3-1M
A 1M token version of the RULER dataset based on the Llama-3 chat template.
It is automatically generated based on the scripts available in the RULER repository: https://github.com/NVIDIA/RULER. It is designed for evaluating the performance of Long Language Models (LLMs) on various tasks with varying sequence lengths.
How to Use
from datasets import load_dataset
LENGTH_IN_STRING = ['4k', '8k', '16k', '32k', '64k', '128k', '256k', '512k', '1M']
TASKS =… See the full description on the dataset page: https://huggingface.co/datasets/self-long/RULER-llama3-1M.ruler-300-seed42
Frozen RULER 300, seed 42
This dataset freezes the exact RULER inputs used by the
short-long-pretraining native evaluation suite.
Repository: bicycleman15/ruler-300-seed42
Rows: 6,300
Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq
Context lengths: 1024, 2048, 4096
Samples per task/length: 300
Seed: 42
Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d
Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.prof_report__runwayml-stable-diffusion-v1-5__multi__24
Dataset Card for "prof_report__runwayml-stable-diffusion-v1-5__multi__24"
More Information needed
georsct
GeoRSCT
A geospatial regression benchmark for evaluating representation–solver compatibility.
GeoRSCT is a benchmark and evaluation framework for studying when geospatial model performance reflects solver quality versus target difficulty, spatial leakage, aggregation effects, scale sensitivity, or representation–solver mismatch.
This release (version 24.0.1) includes 31,789 U.S. ZIP Code Tabulation Areas (ZCTAs), 106 columns spanning 33 ACS features, 37 geospatial enrichment… See the full description on the dataset page: https://huggingface.co/datasets/rudymartin/georsct.rublimp
RuBLiMP
Dataset Description
RuBLiMP, or Russian Benchmark of Linguistic Minimal Pairs, is the first diverse and large-scale benchmark of minimal pairs in Russian.
RuBLiMP includes 45k minimal pairs of sentences that differ in grammaticality and isolate morphological, syntactic, or semantic phenomena. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/rublimp.OlymMATH-eval
OlymMATH Evaluation Results
OlymMATH is a dataset we introduced in Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models by Haoxiang Sun, Yingqian Min, Zhipeng Chen, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, Lei Fang, and Ji-Rong Wen. You can find more information on GitHub and HuggingFace 🤗.
We have made our evaluation results for the avg@{8, 64} and cons@{8, 64} metrics in this dataset publicly available for academic research… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/OlymMATH-eval.RunBugRun-Final
Original Dataset + Tokenized Data + (Buggy + Fixed Embedding Pairs) + Difference Embeddings
Overview
This repository contains 4 related datasets for training a transformation from buggy to fixed code embeddings:
Datasets Included
1. Original Dataset (train-00000-of-00001.parquet)
Description: Legacy RunBugRun Dataset
Format: Parquet file with buggy-fixed code pairs, bug labels, and language
Size: 456,749 samples
Load with:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/RunBugRun-Final.corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.RULER-32768-Qwen2.5-3B-tokenizeriDRAMA-rumble-2024
Dataset Summary
iDRAMA-rumble-2024 is a large-scale dataset of 6,735 podcast videos from Rumble, an alternative Youtube-like platform. Using state-of-the-art models, we extract information across three modalities: 1) text, 2) audio, and 3) video. We detail the methodology for extracting information from podcast videos in the paper and release a first-of-its-kind dataset including data from different modalities:
Metadata: Details about podcast videos, e.g., channel name, video name… See the full description on the dataset page: https://huggingface.co/datasets/iDRAMALab/iDRAMA-rumble-2024.xtac-umi-g1-insert-rubber-stopper
Representative frames from TacVerse's bimanual
demonstrations.
Collected with XTac-UMI-G1 grippers, released as LeRobot
datasets.
This dataset was created using LeRobot.
Explore this dataset with the LeRobot Dataset Viewer.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xtac_umi_g1",
"total_episodes": 10,
"total_frames": 6204,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/xtac-umi-g1-insert-rubber-stopper.AIRBOT_MMK2_storage_rubiks_cube_and_cup
AIRBOT_MMK2_storage_rubiks_cube_and_cup
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_rubiks_cube_and_cup.first_test_run_20260720_125646This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/first_test_run_20260720_125646.earnings-call-data
S&P 500 earnings episodes (2005–2025)
Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/RudrakshNanavaty/earnings-call-data.stihi_ruR1_Lite_move_the_position_of_the_rubiks_cube
R1_Lite_move_the_position_of_the_rubiks_cube
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_rubiks_cube.Public-MakerMods-SO101-runs
bimanual_so101 run visualizer
LeRobot v2.1 playback mirror for robot-3652c537a175cbae.
Original images and telemetry remain in the shared archive. Failed runs are retained; these are not all successful demonstrations.
AIRBOT_MMK2_place_the_umbrella_and_the_ruler
AIRBOT_MMK2_place_the_umbrella_and_the_ruler
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_umbrella_and_the_ruler.Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-1024.statml-arxivhero_run_4_math_codeprism-librispeech-train-100RuHeritage-Corpus
RuHeritage-Corpus
🇬🇧 English Description
RuHeritage-Corpus is a high-quality, curated dataset of Russian classical literature, specifically designed for the pre-training and continued pre-training (CPT) of Large Language Models (LLMs).
The corpus focuses on the Golden and Silver Ages of Russian literature, providing models with exposure to rich vocabulary, complex syntactic structures, and stylistically flawless Russian text, acting as a "quality anchor"… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RuHeritage-Corpus.
