datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sst5
Stanford Sentiment Treebank - Fine-Grained
Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
sst2
Stanford Sentiment Treebank - Binary
Stanford Sentiment Treebank with 2 labels: negative, positive
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
SSTQAsst5
Stanford Sentiment Treebank - Fine-Grained
Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
sst2SST-2-attrpromptThis is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias.
label.txt: the label name for each class
train.jsonl: The original training set.
valid.jsonl: The original validation set.
test.jsonl: The original test set.
simprompt.jsonl: The training data generated by the simple prompt.
attrprompt.jsonl: The training data generated by the attributed prompt.
kyrgyz-sst2
Kyrgyz SST-2
Task
Sentiment Classification (Binary Sentence Classification)
Description
Stanford Sentiment Treebank binary classification task translated from English to Kyrgyz. Each entry contains an English sentence, its Kyrgyz translation, and a binary sentiment label.
Labels: negative, positive
Format: JSONL with fields such as sentence, sentence_ky, and label
Dataset Size
Split
Entries
Train
6,920
Validation
872
Test
1,821… See the full description on the dataset page: https://huggingface.co/datasets/metinovadilet/kyrgyz-sst2.Claude-opus-4.6-TraceInversion-9000x
🌀 Claude-opus-4.6-TraceInversion-9000x
v1.0 Release
A High-Fidelity Reconstructed CoT Dataset via Trace Inversion
📊 9,000 Samples
🧬 Trace Inversion & Negentropy
🛠 SFT & DPO Ready
🔥 Claude 4.6 Distillation
🌐 English & Multilingual
💡 What is Trace Inversion?
In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude) typically hide their internal thinking steps… See the full description on the dataset page: https://huggingface.co/datasets/Sstoryloop725/Claude-opus-4.6-TraceInversion-9000x.Vibe-Coding-Claude-Fable-5sst5-rawsst2-babysst5-bertsst2_coreset_test_new_greeedysst5-bert-scaledsst2_coreset_test_1dua-test-training-datasst2_coreset_test_greedysst2_coreset_test_greedy_reversesst2_coreset_test_greedy_reverse_partitionsst2_coreset_test_greedy_reverse_0.5sst2_coreset_test_greedy_reverse_0.5_normalizedsst2_2_embeddingsmpnet_sst2_embeddingssst2
