datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham.
rte
Glue RTE
This dataset is a port of the official rte dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
AIME_Problem_Set_1983-2024TREC-QC
TREC Question Classification
Question classification in coarse and fine-grained categories.
Source:
Experimental Data for Question Classification
Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002.
diffusion-pretrain-set-ft1
diffusion-pretrain-set-ft1
A multi-source image-caption pretraining dataset assembled from ten upstream
sources via a uniform ingest pipeline. Designed for a full pretrain or finetune
pipeline meant to curate for any major diffusion model preliminary, with the sole
intent to create a more powerful baseline preliminary train and a baseline
for synthesizing images to train the next generation of the VLM model.
This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.Omni-Sets
Omni-Sets
A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities.
590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets
Overview
Config
Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.qqp
Glue QQP
This dataset is a port of the official qqp dataset on the Hub.
Note that the question1 and question2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
SCOPE-OOD-set
SCOPE-60K-OOD: Out-of-Distribution LLM Routing Dataset
Dataset Description
SCOPE-60K-OOD is an out-of-distribution (OOD) evaluation dataset for LLM routing systems. It contains evaluation results from 5 frontier language models that were not seen during training, designed to test the generalization capabilities of routing methods.
Authors
Qi Cao - UC San Diego, PXie Lab
Shuhao Zhang - UC San Diego, PXie Lab
Affiliation
University of California, San… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-OOD-set.mnli
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the matched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
qnli
Glue QNLI
This dataset is a port of the official qnli dataset on the Hub.
Note that the question and sentence columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
mrpc
Glue MRPC
This dataset is a port of the official mrpc dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
leetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.gutenberg_setlebanese_aug_setxglue_nc#xglue nc
This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German.
stsb
Glue STS-B
This dataset is a port of the official sts-b dataset on the Hub.
This is not a classification task, so the label_text column is only included for consistency
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
R1_Lite_tea_service_table_setting
R1_Lite_tea_service_table_setting
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_tea_service_table_setting.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
climbing-holds
[!IMPORTANT]
This dataset is in construction. The current files are raw scans intended for establishing the structure.
Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide.
GUI for contributions
https://setrsoft.github.io/holds-dataset-hub/
Or send your files here
Climbing Holds 3D dataset (SetRsoft)
📋 Project Overview
This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.bing_coronavirus_query_set
Dataset Card for BingCoronavirusQuerySet
Dataset Summary
Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020
example:
load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30")
You can also load the data by country by using queries_by="country".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.hate_speech18fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.cube_picknplace_480x640_set_1wsc_fixed
Glue WSC Fixed
This dataset is a port of the official wsc.fixed dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
so100_grab_ballThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 13031,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setchii/so100_grab_ball.dataset_combine_20251014_setup_1n2_bboxes
dataset_combine_20251014_setup_1n2
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
wnli
Glue WNLI
This dataset is a port of the official wnli dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.wsc
Glue WSC
This dataset is a port of the official wsc dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
