datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham.
rte
Glue RTE
This dataset is a port of the official rte dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
TREC-QC
TREC Question Classification
Question classification in coarse and fine-grained categories.
Source:
Experimental Data for Question Classification
Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002.
qqp
Glue QQP
This dataset is a port of the official qqp dataset on the Hub.
Note that the question1 and question2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
mnli
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the matched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
qnli
Glue QNLI
This dataset is a port of the official qnli dataset on the Hub.
Note that the question and sentence columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
mrpc
Glue MRPC
This dataset is a port of the official mrpc dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
xglue_nc#xglue nc
This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German.
stsb
Glue STS-B
This dataset is a port of the official sts-b dataset on the Hub.
This is not a classification task, so the label_text column is only included for consistency
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
climbing-holds
[!IMPORTANT]
This dataset is in construction. The current files are raw scans intended for establishing the structure.
Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide.
GUI for contributions
https://setrsoft.github.io/holds-dataset-hub/
Or send your files here
Climbing Holds 3D dataset (SetRsoft)
📋 Project Overview
This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.hate_speech18wsc_fixed
Glue WSC Fixed
This dataset is a port of the official wsc.fixed dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
wnli
Glue WNLI
This dataset is a port of the official wnli dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
wsc
Glue WSC
This dataset is a port of the official wsc dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
x402-settlement-census
x402 settlement census — 2026-09-06
We paid 316 conformant x402 hosts as an ordinary buyer and recorded what came back.
Not a survey of what hosts advertise — a record of what they did when real USDC arrived.
outcome
hosts
share
REFUSED
213
67.4%
DELIVERED
100
31.6%
NO_CHALLENGE
2
0.6%
MISMATCH
1
0.3%
Two in three conformant hosts refused a correctly-signed payment. Being listed in a Bazaar index
and answering a valid 402 is not the same as taking money and… See the full description on the dataset page: https://huggingface.co/datasets/csoai/x402-settlement-census.mnli_mm
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the mismatched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.polymarket-settlement-quality-register
Polymarket settlement-quality register
Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features.
Files and viewer
The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations.
Method and source
See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.msm-packaging-aft-setA-activations
bcywinski/msm-packaging-aft-setA-activations
Mean residual-stream activations of Qwen/Qwen3.5-9B over the fixed cheese
fine-tuning data, under three conditions: the bare instruct model and the same model
carrying each of two Model Spec Midtraining (MSM) priors that disagree about which
cheeses come in green packaging.
The point of the set is that the fine-tuning data is identical in all three: these
are the activations of the demonstrations a fine-tune is about to be trained on… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-aft-setA-activations.turkish-over-refusal-set
turkish-over-refusal-set
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set")
An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories.
Finding: guards over-block Turkish, not English
Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.setmm-reach-vs-responsibility
set.mm — Reach vs. Responsibility
Every theorem in a large formal library that reaches an axiom, and whether it actually holds anything up.
This dataset separates two things that look identical from a dependency graph and are not: how many theorems reach an axiom (have some proof-route back to it) versus how much each theorem actually carries (how many other theorems would stop reaching that axiom if it were removed). The headline result across eight major axiom seams of set.mm:… See the full description on the dataset page: https://huggingface.co/datasets/vince-gonzalez/setmm-reach-vs-responsibility.ag_news_training_set_losses
AG News train losses
This dataset is part of an experiment using Rubrix, an open-source Python framework for human-in-the loop NLP data annotation and management.
lima_filteredSet_Evalcses-problem-set-metadataThis dataset contains the metadata for the CSES problem set (e.g. title, time limit, number of test cases, etc).
The data is crawled on December 28, 2024.
Notes:
time_limit is in seconds
memory_limit is in MB
Important: New tasks and categories were added to CSES, see here. This dataset is now outdated and will be updated soon.
llama-3.1-medprm-reward-raw-training-setllama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/jysyoh/llama-3.1-medprm-reward-training-set.aibe-xxi-set-a
AIBE 16–21 MCQ Benchmark
All 100 multiple-choice questions from each of the last six All India Bar Examinations (AIBE XVI–XXI, 2021–2026), as six configs of one dataset, each paired with the most authoritative answer key available and labeled with the same official 19-subject BCI syllabus taxonomy.
Config
Exam
Held
Set
Question source
Answer key
Withdrawn
Multi-answer
aibe16
AIBE XVI
Oct 2021
C
Delhi Law Academy compilation
DLA 4-set key table (no official copy… See the full description on the dataset page: https://huggingface.co/datasets/rushankg/aibe-xxi-set-a.
