CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reasoning-machines /gsm-hard Dataset Summary This is the harder version of gsm8k math reasoning dataset (https://huggingface.co/datasets/gsm8k). We construct this dataset by replacing the numbers in the questions of GSM8K with larger numbers that are less common.  Supported Tasks and Leaderboards This dataset is used to evaluate math reasoning Languages English - Numbers Dataset Structure dataset = load_dataset("reasoning-machines/gsm-hard") DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-machines/gsm-hard.text1K<n<10K66 likes32k downloads4y agoHugging Face02manifold-machines /mmmu-mkimage10K<n<100K0 likes3.1k downloads1y agoHugging Face03aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K2 likes2.8k downloads9mo agoHugging Face04machine-unlearning-bench /data-unlearning-benchDataset for the evaluation of data-unlearning techniques using KLOM (KL-divergence of Margins). How KLOM works: KLOM works by: training N models (original models) Training N fully-retrained models (oracles) on forget set F unlearning forget set F from the original models Comparing the outputs of the unlearned models from the retrained models on different points (specifically, computing the KL divergence between the distribution of margins of oracle models and distribution of… See the full description on the dataset page: https://huggingface.co/datasets/machine-unlearning-bench/data-unlearning-bench.10K<n<100K1 likes2.7k downloads1y agoHugging Face05RGES-PIT /MachineLearning Machine Learning Tier This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection. The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.tabular10B<n<100B1 likes1.7k downloads1mo agoHugging Face06unitreerobotics /G1_WBT_Inspire_Put_Clothes_into_Washing_Machine Data Structure Observations observation.state.ee_state (12) End-effector states of the robot. Computed via forward kinematics (FK) from the root link to the left and right end-effectors. Includes the contribution of the waist. Represented as concatenated poses of both end-effectors. observation.state.hand_state (12 or 2) Finger states for both hands. The dimensionality depends on the hand type. Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Inspire_Put_Clothes_into_Washing_Machine.tabularrobotics100K<n<1M3 likes1.3k downloads6mo agoHugging Face07openzot /machinery zot machinery sessions Every conversation behind every machine in the zot machinery - a software factory where an agent takes the same standing order every half hour, reads the catalogue of what already exists, looks at the last few machines on the shelf, and designs, writes, commissions and ships one brand-new machine: a working control panel for a machine that does not exist, with a live simulation behind it, faults you can trigger and recover from, and the operating manual… See the full description on the dataset page: https://huggingface.co/datasets/openzot/machinery.text-generationn<1K0 likes1.1k downloads2h agoHugging Face08BAAI-DataCube /AgiBotWorld-Beta_G1_task_369_Take_the_clothes_out_of_the_washing_machine agibot_task_369 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 把衣服从洗衣机里拿出来 total_episodes: 1127 total_tasks: 1 size: 87G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_369_Take_the_clothes_out_of_the_washing_machine.videoroboticsn<1K0 likes1.1k downloads9mo agoHugging Face09MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes1.1k downloads9mo agoHugging Face10machine0003 /dataset0 likes905 downloads22d agoHugging Face11Human-Centric-Machine-Learning /tokenization-multiplicity-data Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez. 📂 Dataset Structure The dataset is organized into folders as follows: .\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.text-generation10K<n<100K2 likes885 downloads6mo agoHugging Face12nacloos /machine-playing-paper Paper Supplementary Data Supplementary data for the 30-hour agent runs used in the paper. Layout paper-supplementary/ data/ claude_13agents/ gpt_5agents/ kimi_5agents/ claude_ablated_5agents/ agent01/ hour_01/ recording.clawrec session.jsonl workspace/ ... hour_30/ materials/ initial_workspace/ standard/ claude_ablated/ system_prompts/ Each agentXX is one… See the full description on the dataset page: https://huggingface.co/datasets/nacloos/machine-playing-paper.0 likes764 downloads3mo agoHugging Face13unitreerobotics /G1_WBT_Dex1_Put_Clothes_into_Washing_Machine Data Structure Observations observation.state.ee_state (12) End-effector states of the robot. Computed via forward kinematics (FK) from the root link to the left and right end-effectors. Includes the contribution of the waist. Represented as concatenated poses of both end-effectors. observation.state.hand_state (12 or 2) Finger states for both hands. The dimensionality depends on the hand type. Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Dex1_Put_Clothes_into_Washing_Machine.tabularrobotics100K<n<1M1 likes748 downloads5mo agoHugging Face14text-machine-lab /vocab_filtered_dataset_22B Dataset Card for "vocab_filtered_dataset_22B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES) We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.text100M<n<1B0 likes681 downloads2y agoHugging Face15machine0001 /dataset0 likes654 downloads23d agoHugging Face16pandalla /Machine_Mindset_MBTI_datasetHere are the behavior datasets used for supervised fine-tuning (SFT). And they can also be used for direct preference optimization (DPO). The exact copy can also be found in Github. Prefix 'en' denotes the datasets of the English version. Prefix 'zh' denotes the datasets of the Chinese version. Dataset introduction There are four dimension in MBTI. And there are two opposite attributes within each dimension. To be specific: Energe: Extraversion (E) - Introversion (I)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/Machine_Mindset_MBTI_dataset.text100K<n<1M71 likes550 downloads2y agoHugging Face17unitreerobotics /G1_WBT_Inspire_Put_Clothes_into_Washing_Machine_MainCamOnly Data Structure Observations observation.state.ee_state (12) End-effector states of the robot. Computed via forward kinematics (FK) from the root link to the left and right end-effectors. Includes the contribution of the waist. Represented as concatenated poses of both end-effectors. observation.state.hand_state (12 or 2) Finger states for both hands. The dimensionality depends on the hand type. Inspire Hand (range: 0.0 – 1.0, open → close)… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Inspire_Put_Clothes_into_Washing_Machine_MainCamOnly.tabularrobotics100K<n<1M5 likes547 downloads6mo agoHugging Face18kavehsgh /EuRoC_MAV_Dataset_Machine_Hall_Easy_010 likes528 downloads9mo agoHugging Face19BAAI-DataCube /AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine agibot_task_431 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 用胶囊机煮咖啡 total_episodes: 674 total_tasks: 1 size: 35G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine.videoroboticsn<1K0 likes509 downloads9mo agoHugging Face20RoboCOIN /R1_Lite_take_clothes_out_of_the_washing_machinegated R1_Lite_take_clothes_out_of_the_washing_machine 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_take_clothes_out_of_the_washing_machine.tabularrobotics100K<n<1M0 likes501 downloads9mo agoHugging Face21christian0420 /drifting-vla-v2-g1_inspire_washing_machine_full0 likes458 downloads6mo agoHugging Face22unitreerobotics /G1_WBT_Brainco_Put_Clothes_Into_Washing_Machinevideon<1K1 likes434 downloads8d agoHugging Face23cahlen /ramanujan-machine-results Ramanujan Machine — GPU Formula Discovery Results GPU-accelerated search for new continued fraction formulas for mathematical constants, inspired by Raayoni et al. (2024). Part of the bigcompute.science project. AI-audited, not peer-reviewed. Key Findings (Updated 2026-04-07) 586 billion equal-degree polynomial CFs exhausted (v1 kernel, degrees 1-8) — zero new transcendental formulas discovered. 7,030 "transcendental hits" were double-precision false positives —… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/ramanujan-machine-results.0 likes423 downloads6mo agoHugging Face24BAAI-DataCube /AgiBotWorld-Beta_G1_task_465_Wash_clothes_in_a_washing_machine agibot_task_465 This dataset converts the AgiBot format uniformly into LeRobot V3.0. Dataset Statistics robot_name: G1 end_effector: 夹爪 task: 用洗衣机洗衣服。 total_episodes: 615 total_tasks: 1 size: 69G Dataset Structure ├── data │ └── chunk-xxx │ ├── file-xxx.parquet ├── meta │ ├── episodes │ │ └── chunk-xxx │ │ └── file-xxx.parquet │ ├── info.json │ ├── stats.json │ └── tasks.parquet └── videos ├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_465_Wash_clothes_in_a_washing_machine.videoroboticsn<1K0 likes417 downloads9mo agoHugging Face25pgurazada1 /machine-failure-mlops-demo-logstabular1K<n<10K0 likes354 downloads2y agoHugging Face26DigitalUmuganda /monolingual_machine_translation_datatext100K<n<1M0 likes337 downloads3y agoHugging Face27CohereLabs /dolly-machine-translated-v2 Dolly Machine Translated (v2) Dataset Description Dolly Machine Translated (v2) is a multilingual evaluation-only release built from a curated subset of Databricks Dolly 15k prompts. It contains the original English prompts plus machine translations in 66 non-English languages, with the English source prompts included as the en config for reference. Each language is provided as a separate config (subset). All language codes use ISO 639-1 two-letter codes. Each row carries… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/dolly-machine-translated-v2.text10K<n<100K2 likes292 downloads5mo agoHugging Face28text-machine-lab /vocab_filtered_dataset_2.1B Dataset Card for "vocab_filtered_dataset_2.1B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES) We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_2.1B.text1M<n<10M0 likes268 downloads2y agoHugging Face29hopeahilton /toward-life-machine-readableHuman-facing blog page: http://noharmscripture.com/ license: mit task_categories: - question-answering - text-generation language: - en tags: - theology - harm-reduction - biblical-studies - safety - religious-qa - crisis-intervention - pastoral-care - ai-safety - translation-literacy - liberation-theology - wesleyan pretty_name: "Toward Life: Biblical Harm Reduction Index" size_categories: - n<1K Toward Life: Biblical Harm Reduction Index… See the full description on the dataset page: https://huggingface.co/datasets/hopeahilton/toward-life-machine-readable.0 likes259 downloads8mo agoHugging Face30DigitalUmuganda /kinyarwanda-english-machine-translation-dataset Kinyarwanda English Parallel Datasets for Machine translation A 48,000 Kinyarwanda English Parallel datasets for machine translation, made by curating and translating normal Kinyarwanda sentences into English 4 likes250 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.