datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task363_sst2_polarity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task363_sst2_polarity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task363_sst2_polarity_classification.Claude-opus-4.6-TraceInversion-9000x
🌀 Claude-opus-4.6-TraceInversion-9000x
v1.0 Release
A High-Fidelity Reconstructed CoT Dataset via Trace Inversion
📊 9,000 Samples
🧬 Trace Inversion & Negentropy
🛠 SFT & DPO Ready
🔥 Claude 4.6 Distillation
🌐 English & Multilingual
💡 What is Trace Inversion?
In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude) typically hide their internal thinking steps… See the full description on the dataset page: https://huggingface.co/datasets/Sstoryloop725/Claude-opus-4.6-TraceInversion-9000x.sst_pt-pt
Simple Safety Tests-PT
Portuguese machine translation of Simple Safety Tests, a benchmark for evaluating model safety and harmful content detection.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/Bertievidgen/SimpleSafetyTests
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/sst_pt-pt.
