datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
feedback-prize-english-language-learning-fluencymultilingual-fluencytask138_detoxifying-lms_classification_fluency
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task138_detoxifying-lms_classification_fluency
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task138_detoxifying-lms_classification_fluency.normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Citation
@misc{samuel2025fluentalignmentdisfluentjudges,
title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages},
author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov},
year={2025},
eprint={2512.08777},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.fluency-pairs-gec-onlysep28k-fluencybank-stutter-datasetAdaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_AnimalsThe contents of this dataset are from the Forager computer program. Its interactive interface can be found at https://forager.research.bowdoin.edu.
Please read the LICENSE file in this data set and check for further information on the Forager GitHub repository (https://github.com/thelexiconlab/forager-web)
in addition to the docs tab of the web interface linked above.
Citation:
Kumar, A.A., Apsel, M., Zhang, L., Xing, N., Jones. M.N. (2023). forager: A Python package and web interface for… See the full description on the dataset page: https://huggingface.co/datasets/natebranda/Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_Animals.speechocean762_fluencyfluency-edits-geminifluency-pairs-gec-onklycompletion-coherence-fluencyfluency-edits-gec-onlyfluency-agumented-editskeural-v2-fluency-v2
Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition |
Methodology |
License |
Limitations
Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.fluency-agumented-binaryfluency-pairsfluency-editsL2EnglishScoring_speechocean762_fluencySoofi-German-Fluency-DPO
Dataset Card — German Fluency Preference Dataset
Overview
This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus.
All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted.
Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-German-Fluency-DPO.fluency-pairs-gec-only-single-editfluency-datasetfluency-edits-gec-only-extended3sec_stuttering_only_fluencybankcoedit-fluencytrain-fluency-datasetfluency-agumented-pairsL2EnglishScoring_speechocean762_fluency_v2L2EnglishScoring_speechocean762_fluencyspeechocean762_fluency_4_training3s_fluencybank
