datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sst_migaretThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 61,
"total_frames": 53650,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/daxum34/sst_migaret.SSTQAglue_augmented_sst2
Dataset Card for glue_augmented_sst2
Dataset Description
Augmented SST-2 dataset
Reference: https://huggingface.co/datasets/glue
kyrgyz-sst2
Kyrgyz SST-2
Task
Sentiment Classification (Binary Sentence Classification)
Description
Stanford Sentiment Treebank binary classification task translated from English to Kyrgyz. Each entry contains an English sentence, its Kyrgyz translation, and a binary sentiment label.
Labels: negative, positive
Format: JSONL with fields such as sentence, sentence_ky, and label
Dataset Size
Split
Entries
Train
6,920
Validation
872
Test
1,821… See the full description on the dataset page: https://huggingface.co/datasets/metinovadilet/kyrgyz-sst2.UDR_SST-5
Dataset Card for "UDR_SST-5"
More Information needed
probe-robustness-sst2
sst2
Dataset repo: wrynx/probe-robustness-sst2
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 68221
Records per split:
test: 10910
train: 46918
valid: 10393
Number of classes: 2
Records per class:
0: 30208
1: 38013
Records per class per split:
test:
0: 4900
1: 6010
train:
0: 20712
1: 26206
valid:
0: 4596
1: 5797
Original dataset README (from… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-sst2.sst2-es-mt
STT-2 Spanish
A Spanish translation (using EasyNMT) of the SST-2 Dataset
For more information check the official Model Card
sst2sst2_Full-p_05sst2_Sampled-p_1SST_sentiment_fairness_data
Sentiment fairness dataset
================================
This dataset is to measure gender fairness in the downstream task of sentiment analysis. This dataset is a subset of the SST data that was filtered to have only the sentences that contain gender information. The python code used to create this dataset can be found in the prepare_sst.ipyth file.
Then the filtered datset was labeled by 4 human annotators who are the authors of this dataset. The annotations… See the full description on the dataset page: https://huggingface.co/datasets/fatmaElsafoury2022/SST_sentiment_fairness_data.sst2_Full-p_1UDR_SST-2
Dataset Card for "UDR_SST-2"
More Information needed
dataset_env_SST_SP6_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesianThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 1698,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cbrian/dataset_env_SST_SP6_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesian.sstllm-metric-glue-sst2dataset_env_SST_SP8_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesianThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 1825,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cbrian/dataset_env_SST_SP8_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesian.populism-llm
Populism-LLM
A cross-validated LLM annotation of European party manifestos on
populism and liberalism dimensions. Three independent large language models
— Claude Sonnet 4.6 (Anthropic), GPT-4.1-mini (OpenAI), and Gemini Flash
(Google) — score every CMP manifesto using an identical strict JSON
schema. Each score is backed by a verbatim quote and contextual snippet.
🚧 Preview / draft release. A peer-reviewed paper-companion v1.0
release with DOI is forthcoming. Until then… See the full description on the dataset page: https://huggingface.co/datasets/sstoeckl/populism-llm.cobie_sst2
Dataset Card for cobie_sst2
This dataset is a modification of the original SST-2 dataset for LLM cognitive bias evaluation.
Language(s)
English (en)
Dataset Summary
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
The corpus is based on the dataset introduced by Pang and Lee (2005) and consists of 11,855 single sentences extracted from movie… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/cobie_sst2.sst-2sst2_balancedMULTI_VALUE_sst2_drop_inf_to
Dataset Card for "MULTI_VALUE_sst2_drop_inf_to"
More Information needed
sst2_Sampled-p_05MULTI_VALUE_sst2_double_modals
Dataset Card for "MULTI_VALUE_sst2_double_modals"
More Information needed
sst2_mnli_qqp_llama1b_modified
Multi-Task Dataset: SST-2 + MNLI + QQP (Modified for LLaMA 1B)
This dataset is a combination of SST-2, MNLI, and QQP for multi-task learning.
It is preprocessed and tokenized specifically for training with the LLaMA-1B model.
Modifications:
Each example includes a task prefix:
SST-2: "Task: SST2 | Sentence: ..."
MNLI: "Task: MNLI | Premise: ... Hypothesis: ..."
QQP: "Task: QQP | Q1: ... Q2: ..."
Labels are standardized to integer format.
Tokenized using the LLaMA-1B… See the full description on the dataset page: https://huggingface.co/datasets/emirhanboge/sst2_mnli_qqp_llama1b_modified.dataset_env_SST_SP5_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesianThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 1591,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cbrian/dataset_env_SST_SP5_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesian.sst2-poisoned-target-1-testsetsst2-data_augmentationsst2-ka
sst2-ka
Georgian translation of the SST-2 (Stanford Sentiment Treebank) benchmark.
Dataset Summary
Property
Value
Examples
481
Splits
validation
Languages
Georgian, English
Task
Sentiment Classification
Data Fields
sentence: Sentence (English)
label: Sentiment label (0=negative, 1=positive)
idx: Example index
sentence_ka: Sentence (Georgian)
Translation Methodology
Translation model generates initial Georgian translation… See the full description on the dataset page: https://huggingface.co/datasets/tbilisi-ai-lab/sst2-ka.dataset_env_SST_SP2_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesianThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 1825,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cbrian/dataset_env_SST_SP2_WC1_TC1_task_pickplaceblueblock_numepi_10_ctrl_cartesian.
