CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tcapelle /feedback-prize-english-language-learning-fluencytabular1K<n<10K1 likes100 downloads2y agoHugging Face02geoalgo /multilingual-fluencytext1K<n<10K0 likes85 downloads4mo agoHugging Face03Lots-of-LoRAs /task138_detoxifying-lms_classification_fluency Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task138_detoxifying-lms_classification_fluency Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task138_detoxifying-lms_classification_fluency.texttext-generationn<1K0 likes84 downloads2y agoHugging Face04ltg /normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages Citation @misc{samuel2025fluentalignmentdisfluentjudges, title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages}, author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov}, year={2025}, eprint={2512.08777}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.textn<1K0 likes78 downloads10mo agoHugging Face05timonziegenbein /fluency-pairs-gec-onlytext10K<n<100K0 likes72 downloads11mo agoHugging Face06saeedzou /sep28k-fluencybank-stutter-datasetaudio10K<n<100K0 likes58 downloads2mo agoHugging Face07natebranda /Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_AnimalsThe contents of this dataset are from the Forager computer program. Its interactive interface can be found at https://forager.research.bowdoin.edu. Please read the LICENSE file in this data set and check for further information on the Forager GitHub repository (https://github.com/thelexiconlab/forager-web) in addition to the docs tab of the web interface linked above. Citation: Kumar, A.A., Apsel, M., Zhang, L., Xing, N., Jones. M.N. (2023). forager: A Python package and web interface for… See the full description on the dataset page: https://huggingface.co/datasets/natebranda/Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_Animals.text0 likes53 downloads2mo agoHugging Face08papasega /speechocean762_fluencyaudio1K<n<10K1 likes34 downloads3y agoHugging Face09timonziegenbein /fluency-edits-geminitext1K<n<10K0 likes33 downloads11mo agoHugging Face10timonziegenbein /fluency-pairs-gec-onklytext10K<n<100K0 likes29 downloads11mo agoHugging Face11tcapelle /completion-coherence-fluencytabular10K<n<100K1 likes25 downloads2y agoHugging Face12timonziegenbein /fluency-edits-gec-onlytext1K<n<10K0 likes25 downloads11mo agoHugging Face13timonziegenbein /fluency-agumented-editstext100K<n<1M0 likes24 downloads11mo agoHugging Face14mkd-minju /keural-v2-fluency-v2 Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning Introduction | Dataset Composition | Methodology | License | Limitations Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.texttext-generation100K<n<1M0 likes21 downloads1mo agoHugging Face15timonziegenbein /fluency-agumented-binarytext100K<n<1M0 likes20 downloads1y agoHugging Face16timonziegenbein /fluency-pairstext10K<n<100K0 likes19 downloads11mo agoHugging Face17timonziegenbein /fluency-editstext10K<n<100K0 likes19 downloads11mo agoHugging Face18speech31 /L2EnglishScoring_speechocean762_fluencyaudion<1K0 likes18 downloads2y agoHugging Face19toroe /Soofi-German-Fluency-DPO Dataset Card — German Fluency Preference Dataset Overview This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus. All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted. Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-German-Fluency-DPO.texttext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face20timonziegenbein /fluency-pairs-gec-only-single-edittext1K<n<10K0 likes16 downloads11mo agoHugging Face21tcapelle /fluency-datasettext10K<n<100K0 likes14 downloads2y agoHugging Face22timonziegenbein /fluency-edits-gec-only-extendedtext10K<n<100K0 likes13 downloads11mo agoHugging Face23khushinakra /3sec_stuttering_only_fluencybankaudion<1K0 likes11 downloads2y agoHugging Face24tcapelle /coedit-fluencytext10K<n<100K0 likes10 downloads2y agoHugging Face25tcapelle /train-fluency-datasettext10K<n<100K0 likes9 downloads2y agoHugging Face26timonziegenbein /fluency-agumented-pairstext100K<n<1M0 likes9 downloads11mo agoHugging Face27speech31 /L2EnglishScoring_speechocean762_fluency_v2audion<1K0 likes8 downloads2y agoHugging Face28DynamicSuperb /L2EnglishScoring_speechocean762_fluencyaudion<1K0 likes8 downloads2y agoHugging Face29papasega /speechocean762_fluency_4_trainingaudio1K<n<10K1 likes6 downloads3y agoHugging Face30khushinakra /3s_fluencybankaudio1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.