CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes6.3k downloads5y agoHugging Face02SetFit /rte Glue RTE This dataset is a port of the official rte dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K2 likes4.8k downloads5y agoHugging Face03Prompt48 /AIME_Problem_Set_1983-2024tabularn<1K0 likes4.2k downloads2y agoHugging Face04SetFit /TREC-QC TREC Question Classification Question classification in coarse and fine-grained categories. Source: Experimental Data for Question Classification Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002. tabular1K<n<10K0 likes3.3k downloads5y agoHugging Face05AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.9k downloads3mo agoHugging Face06TrustAIRLab /forbidden_question_set Forbidden Question Set This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy. We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.tabularn<1K7 likes1.7k downloads2y agoHugging Face07MBZUAI /Omni-Setsgated Omni-Sets A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities. 590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets Overview Config Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.audio100K<n<1M0 likes1.6k downloads22d agoHugging Face08SetFit /qqp Glue QQP This dataset is a port of the official qqp dataset on the Hub. Note that the question1 and question2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M6 likes1.6k downloads5y agoHugging Face09Cooolder /SCOPE-OOD-set SCOPE-60K-OOD: Out-of-Distribution LLM Routing Dataset Dataset Description SCOPE-60K-OOD is an out-of-distribution (OOD) evaluation dataset for LLM routing systems. It contains evaluation results from 5 frontier language models that were not seen during training, designed to test the generalization capabilities of routing methods. Authors Qi Cao - UC San Diego, PXie Lab Shuhao Zhang - UC San Diego, PXie Lab Affiliation University of California, San… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-OOD-set.tabulartext-classification1K<n<10K0 likes1.6k downloads8mo agoHugging Face10SetFit /mnli Glue MNLI This dataset is a port of the official mnli dataset on the Hub. It contains the matched version. Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M8 likes1.5k downloads5y agoHugging Face11SetFit /qnli Glue QNLI This dataset is a port of the official qnli dataset on the Hub. Note that the question and sentence columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M2 likes1.3k downloads5y agoHugging Face12SetFit /mrpc Glue MRPC This dataset is a port of the official mrpc dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K17 likes1.2k downloads5y agoHugging Face13kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K9 likes900 downloads1y agoHugging Face14Pclanglais /gutenberg_settabular1M<n<10M0 likes717 downloads2y agoHugging Face15WissMah /lebanese_aug_setimage10K<n<100K0 likes672 downloads11mo agoHugging Face16SetFit /xglue_nc#xglue nc This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German. tabular100K<n<1M0 likes662 downloads2y agoHugging Face17SetFit /stsb Glue STS-B This dataset is a port of the official sts-b dataset on the Hub. This is not a classification task, so the label_text column is only included for consistency Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K1 likes648 downloads5y agoHugging Face18RoboCOIN /R1_Lite_tea_service_table_settinggated R1_Lite_tea_service_table_setting 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_tea_service_table_setting.tabularrobotics100K<n<1M0 likes633 downloads9mo agoHugging Face19SetFit /go_emotions GoEmotions This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification. tabular10K<n<100K13 likes616 downloads4y agoHugging Face20setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K0 likes580 downloads5mo agoHugging Face21microsoft /bing_coronavirus_query_set Dataset Card for BingCoronavirusQuerySet Dataset Summary Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020 example: load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30") You can also load the data by country by using queries_by="country". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.tabulartext-classification100K<n<1M1 likes455 downloads3y agoHugging Face22SetFit /hate_speech18tabular10K<n<100K3 likes444 downloads5y agoHugging Face23physicl-community /fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde Fast-Food Cleaning Robot — Floor Mess Dataset Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.imagen<1K0 likes437 downloads21d agoHugging Face24DreamMachines /cube_picknplace_480x640_set_1tabular100K<n<1M0 likes344 downloads4mo agoHugging Face25SetFit /wsc_fixed Glue WSC Fixed This dataset is a port of the official wsc.fixed dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K1 likes311 downloads4y agoHugging Face26Setchii /so100_grab_ballThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13031, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setchii/so100_grab_ball.tabularrobotics10K<n<100K0 likes293 downloads1y agoHugging Face27phospho-app /dataset_combine_20251014_setup_1n2_bboxes dataset_combine_20251014_setup_1n2 This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularrobotics100K<n<1M0 likes271 downloads11mo agoHugging Face28SetFit /wnli Glue WNLI This dataset is a port of the official wnli dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabularn<1K0 likes250 downloads5y agoHugging Face29NoeFlandre /landuse-sentence-relevance-golden-human-set Land-use sentence relevance golden human set This release contains the final 300-row V3 benchmark in English plus one parallel CSV for each of the 84 non-English project-provided sat-3l-sm language codes. There are 85 language files in total. Files Every file is at data/translations/<iso>/v3-final-<iso>.csv. The nine columns are: sentence, label, polygon_name, h3_cell, latitude, longitude, source, region, source_url. The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.tabulartext-classification10K<n<100K0 likes245 downloads3d agoHugging Face30SetFit /wsc Glue WSC This dataset is a port of the official wsc dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K0 likes243 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.