datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsm-hard
Dataset Summary
This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k).
We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.
Supported Tasks and Leaderboards
This dataset is used to evaluate math reasoning
Languages
English - Numbers
Dataset Structure
dataset = load_dataset("reasoning-machines/gsm-hard")
DatasetDict({
train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.mmmu-mkNLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.data-unlearning-benchDataset for the evaluation of data-unlearning techniques using KLOM (KL-divergence of Margins).
How KLOM works:
KLOM works by:
training N models (original models)
Training N fully-retrained models (oracles) on forget set F
unlearning forget set F from the original models
Comparing the outputs of the unlearned models from the retrained models on different points
(specifically, computing the KL divergence between the distribution of margins of oracle models and distribution of… See the full description on the dataset page: https://huggingface.co/datasets/machine-unlearning-bench/data-unlearning-bench.MachineLearning
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.G1_WBT_Inspire_Put_Clothes_into_Washing_Machine
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Inspire_Put_Clothes_into_Washing_Machine.machinery
zot machinery sessions
Every conversation behind every machine in the
zot machinery - a software factory
where an agent takes the same standing order every half hour, reads the
catalogue of what already exists, looks at the last few machines on the shelf,
and designs, writes, commissions and ships one brand-new machine: a working
control panel for a machine that does not exist, with a live simulation behind
it, faults you can trigger and recover from, and the operating manual… See the full description on the dataset page: https://huggingface.co/datasets/openzot/machinery.AgiBotWorld-Beta_G1_task_369_Take_the_clothes_out_of_the_washing_machine
agibot_task_369
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 把衣服从洗衣机里拿出来
total_episodes: 1127
total_tasks: 1
size: 87G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_369_Take_the_clothes_out_of_the_washing_machine.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.datasettokenization-multiplicity-data
Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez.
📂 Dataset Structure
The dataset is organized into folders as follows:
.\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl
where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.machine-playing-paper
Paper Supplementary Data
Supplementary data for the 30-hour agent runs used in the paper.
Layout
paper-supplementary/
data/
claude_13agents/
gpt_5agents/
kimi_5agents/
claude_ablated_5agents/
agent01/
hour_01/
recording.clawrec
session.jsonl
workspace/
...
hour_30/
materials/
initial_workspace/
standard/
claude_ablated/
system_prompts/
Each agentXX is one… See the full description on the dataset page: https://huggingface.co/datasets/nacloos/machine-playing-paper.G1_WBT_Dex1_Put_Clothes_into_Washing_Machine
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Dex1_Put_Clothes_into_Washing_Machine.vocab_filtered_dataset_22B
Dataset Card for "vocab_filtered_dataset_22B"
Dataset Summary
This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES)
We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.datasetMachine_Mindset_MBTI_datasetHere are the behavior datasets used for supervised fine-tuning (SFT). And they can also be used for direct preference optimization (DPO).
The exact copy can also be found in Github.
Prefix 'en' denotes the datasets of the English version.
Prefix 'zh' denotes the datasets of the Chinese version.
Dataset introduction
There are four dimension in MBTI. And there are two opposite attributes within each dimension.
To be specific:
Energe: Extraversion (E) - Introversion (I)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/Machine_Mindset_MBTI_dataset.G1_WBT_Inspire_Put_Clothes_into_Washing_Machine_MainCamOnly
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Inspire_Put_Clothes_into_Washing_Machine_MainCamOnly.EuRoC_MAV_Dataset_Machine_Hall_Easy_01AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine
agibot_task_431
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 用胶囊机煮咖啡
total_episodes: 674
total_tasks: 1
size: 35G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine.R1_Lite_take_clothes_out_of_the_washing_machine
R1_Lite_take_clothes_out_of_the_washing_machine
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_take_clothes_out_of_the_washing_machine.drifting-vla-v2-g1_inspire_washing_machine_fullG1_WBT_Brainco_Put_Clothes_Into_Washing_Machineramanujan-machine-results
Ramanujan Machine — GPU Formula Discovery Results
GPU-accelerated search for new continued fraction formulas for mathematical constants, inspired by Raayoni et al. (2024).
Part of the bigcompute.science project. AI-audited, not peer-reviewed.
Key Findings (Updated 2026-04-07)
586 billion equal-degree polynomial CFs exhausted (v1 kernel, degrees 1-8) — zero new transcendental formulas discovered.
7,030 "transcendental hits" were double-precision false positives —… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/ramanujan-machine-results.AgiBotWorld-Beta_G1_task_465_Wash_clothes_in_a_washing_machine
agibot_task_465
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 用洗衣机洗衣服。
total_episodes: 615
total_tasks: 1
size: 69G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_465_Wash_clothes_in_a_washing_machine.machine-failure-mlops-demo-logsmonolingual_machine_translation_datadolly-machine-translated-v2
Dolly Machine Translated (v2)
Dataset Description
Dolly Machine Translated (v2) is a multilingual evaluation-only release built from a curated subset of Databricks Dolly 15k prompts. It contains the original English prompts plus machine translations in 66 non-English languages, with the English source prompts included as the en config for reference.
Each language is provided as a separate config (subset). All language codes use ISO 639-1 two-letter codes. Each row carries… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/dolly-machine-translated-v2.vocab_filtered_dataset_2.1B
Dataset Card for "vocab_filtered_dataset_2.1B"
Dataset Summary
This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES)
We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_2.1B.toward-life-machine-readableHuman-facing blog page: http://noharmscripture.com/
license: mit
task_categories:
- question-answering
- text-generation
language:
- en
tags:
- theology
- harm-reduction
- biblical-studies
- safety
- religious-qa
- crisis-intervention
- pastoral-care
- ai-safety
- translation-literacy
- liberation-theology
- wesleyan
pretty_name: "Toward Life: Biblical Harm Reduction Index"
size_categories:
- n<1K
Toward Life: Biblical Harm Reduction Index… See the full description on the dataset page: https://huggingface.co/datasets/hopeahilton/toward-life-machine-readable.kinyarwanda-english-machine-translation-dataset
Kinyarwanda English Parallel Datasets for Machine translation
A 48,000 Kinyarwanda English Parallel datasets for machine translation, made by curating and translating normal Kinyarwanda sentences into English
