CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes6.3k downloads5y agoHugging Face02SetFit /rte Glue RTE This dataset is a port of the official rte dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K2 likes4.8k downloads5y agoHugging Face03SetFit /TREC-QC TREC Question Classification Question classification in coarse and fine-grained categories. Source: Experimental Data for Question Classification Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002. tabular1K<n<10K0 likes3.3k downloads5y agoHugging Face04SetFit /qqp Glue QQP This dataset is a port of the official qqp dataset on the Hub. Note that the question1 and question2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M6 likes1.6k downloads5y agoHugging Face05SetFit /mnli Glue MNLI This dataset is a port of the official mnli dataset on the Hub. It contains the matched version. Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M8 likes1.5k downloads5y agoHugging Face06SetFit /qnli Glue QNLI This dataset is a port of the official qnli dataset on the Hub. Note that the question and sentence columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M2 likes1.3k downloads5y agoHugging Face07SetFit /mrpc Glue MRPC This dataset is a port of the official mrpc dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K17 likes1.2k downloads5y agoHugging Face08SetFit /xglue_nc#xglue nc This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German. tabular100K<n<1M0 likes662 downloads2y agoHugging Face09SetFit /stsb Glue STS-B This dataset is a port of the official sts-b dataset on the Hub. This is not a classification task, so the label_text column is only included for consistency Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K1 likes648 downloads5y agoHugging Face10SetFit /go_emotions GoEmotions This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification. tabular10K<n<100K13 likes616 downloads4y agoHugging Face11setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K0 likes580 downloads5mo agoHugging Face12SetFit /hate_speech18tabular10K<n<100K3 likes444 downloads5y agoHugging Face13SetFit /wsc_fixed Glue WSC Fixed This dataset is a port of the official wsc.fixed dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K1 likes311 downloads4y agoHugging Face14SetFit /wnli Glue WNLI This dataset is a port of the official wnli dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabularn<1K0 likes250 downloads5y agoHugging Face15SetFit /wsc Glue WSC This dataset is a port of the official wsc dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K0 likes243 downloads4y agoHugging Face16csoai /x402-settlement-census x402 settlement census — 2026-09-06 We paid 316 conformant x402 hosts as an ordinary buyer and recorded what came back. Not a survey of what hosts advertise — a record of what they did when real USDC arrived. outcome hosts share REFUSED 213 67.4% DELIVERED 100 31.6% NO_CHALLENGE 2 0.6% MISMATCH 1 0.3% Two in three conformant hosts refused a correctly-signed payment. Being listed in a Bazaar index and answering a valid 402 is not the same as taking money and… See the full description on the dataset page: https://huggingface.co/datasets/csoai/x402-settlement-census.tabularother1K<n<10K1 likes172 downloads10h agoHugging Face17SetFit /mnli_mm Glue MNLI This dataset is a port of the official mnli dataset on the Hub. It contains the mismatched version. Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M0 likes171 downloads5y agoHugging Face18dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes136 downloads1y agoHugging Face19playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face20ailinsun /polymarket-settlement-quality-register Polymarket settlement-quality register Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features. Files and viewer The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations. Method and source See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.tabularn<1K0 likes124 downloads12d agoHugging Face21bcywinski /msm-packaging-aft-setA-activations bcywinski/msm-packaging-aft-setA-activations Mean residual-stream activations of Qwen/Qwen3.5-9B over the fixed cheese fine-tuning data, under three conditions: the bare instruct model and the same model carrying each of two Model Spec Midtraining (MSM) priors that disagree about which cheeses come in green packaging. The point of the set is that the fine-tuning data is identical in all three: these are the activations of the demonstrations a fine-tune is about to be trained on… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-aft-setA-activations.tabular1K<n<10K0 likes99 downloads15d agoHugging Face22fevziegeyurtsevenler /turkish-over-refusal-set turkish-over-refusal-set from datasets import load_dataset ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set") An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories. Finding: guards over-block Turkish, not English Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.tabulartext-classificationn<1K0 likes97 downloads2mo agoHugging Face23vince-gonzalez /setmm-reach-vs-responsibility set.mm — Reach vs. Responsibility Every theorem in a large formal library that reaches an axiom, and whether it actually holds anything up. This dataset separates two things that look identical from a dependency graph and are not: how many theorems reach an axiom (have some proof-route back to it) versus how much each theorem actually carries (how many other theorems would stop reaching that axiom if it were removed). The headline result across eight major axiom seams of set.mm:… See the full description on the dataset page: https://huggingface.co/datasets/vince-gonzalez/setmm-reach-vs-responsibility.tabulargraph-ml100K<n<1M0 likes69 downloads29d agoHugging Face24dvilasuero /ag_news_training_set_losses AG News train losses This dataset is part of an experiment using Rubrix, an open-source Python framework for human-in-the loop NLP data annotation and management. tabular100K<n<1M0 likes44 downloads5y agoHugging Face25SetonLiang2 /lima_filteredtabularn<1K0 likes40 downloads5mo agoHugging Face26AarushSah /Set_Evaltabular1K<n<10K1 likes34 downloads2y agoHugging Face27minhnguyent546 /cses-problem-set-metadataThis dataset contains the metadata for the CSES problem set (e.g. title, time limit, number of test cases, etc). The data is crawled on December 28, 2024. Notes: time_limit is in seconds memory_limit is in MB Important: New tasks and categories were added to CSES, see here. This dataset is now outdated and will be updated soon. tabularn<1K0 likes29 downloads1y agoHugging Face28dmis-lab /llama-3.1-medprm-reward-raw-training-settabular10K<n<100K0 likes25 downloads1y agoHugging Face29jysyoh /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/jysyoh/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face30rushankg /aibe-xxi-set-a AIBE 16–21 MCQ Benchmark All 100 multiple-choice questions from each of the last six All India Bar Examinations (AIBE XVI–XXI, 2021–2026), as six configs of one dataset, each paired with the most authoritative answer key available and labeled with the same official 19-subject BCI syllabus taxonomy. Config Exam Held Set Question source Answer key Withdrawn Multi-answer aibe16 AIBE XVI Oct 2021 C Delhi Law Academy compilation DLA 4-set key table (no official copy… See the full description on the dataset page: https://huggingface.co/datasets/rushankg/aibe-xxi-set-a.documentquestion-answeringn<1K0 likes15 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.