CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamzhangship /fluency-ecdict-offline Fluency ECDICT Offline Pack This repository hosts the versioned ECDICT SQLite packages downloaded by the Fluency app for offline video word lookup and local subtitle focus-word matching. Current release Field Value Data version ecdict-bc015ed2-focus13-v2 Schema 2 Entries 770,611 Focus-word lists 13 ZIP size 54,594,964 bytes SQLite size 131,756,032 bytes ZIP SHA-256 1e745ea698878772226a7df584129409dd4b27cb7ea26b4259bc770e7533352f Source… See the full description on the dataset page: https://huggingface.co/datasets/iamzhangship/fluency-ecdict-offline.100K<n<1M0 likes500 downloads2mo agoHugging Face02Carson-Shively /fluency-trainer0 likes136 downloads2mo agoHugging Face03tcapelle /feedback-prize-english-language-learning-fluencytabular1K<n<10K1 likes99 downloads2y agoHugging Face04geoalgo /multilingual-fluencytext1K<n<10K0 likes89 downloads4mo agoHugging Face05Lots-of-LoRAs /task138_detoxifying-lms_classification_fluency Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task138_detoxifying-lms_classification_fluency Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task138_detoxifying-lms_classification_fluency.texttext-generationn<1K0 likes81 downloads2y agoHugging Face06timonziegenbein /fluency-pairs-gec-onlytext10K<n<100K0 likes70 downloads11mo agoHugging Face07ltg /normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages Citation @misc{samuel2025fluentalignmentdisfluentjudges, title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages}, author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov}, year={2025}, eprint={2512.08777}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.textn<1K0 likes64 downloads10mo agoHugging Face08saeedzou /sep28k-fluencybank-stutter-datasetaudio10K<n<100K0 likes56 downloads2mo agoHugging Face09natebranda /Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_AnimalsThe contents of this dataset are from the Forager computer program. Its interactive interface can be found at https://forager.research.bowdoin.edu. Please read the LICENSE file in this data set and check for further information on the Forager GitHub repository (https://github.com/thelexiconlab/forager-web) in addition to the docs tab of the web interface linked above. Citation: Kumar, A.A., Apsel, M., Zhang, L., Xing, N., Jones. M.N. (2023). forager: A Python package and web interface for… See the full description on the dataset page: https://huggingface.co/datasets/natebranda/Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_Animals.text0 likes52 downloads2mo agoHugging Face10papasega /speechocean762_fluencyaudio1K<n<10K1 likes35 downloads3y agoHugging Face11timonziegenbein /fluency-edits-geminitext1K<n<10K0 likes33 downloads11mo agoHugging Face12timonziegenbein /fluency-pairs-gec-onklytext10K<n<100K0 likes28 downloads11mo agoHugging Face13timonziegenbein /fluency-edits-gec-onlytext1K<n<10K0 likes24 downloads11mo agoHugging Face14mkd-minju /keural-v2-fluency Keural-v2 Fluency: A Curated Korean–English Conversational Corpus for LLM Fine-Tuning Introduction | Dataset Composition (Revisions) | Methodology | License | Limitations Status: Private staging — not yet cleared for public release (pending §4 evaluation and second-party license audit). 1. Introduction Keural-v2 Fluency is the "Area 1" component of the Keural-v2 Korean SFT training corpus, purpose-built for DeepSeek-V4-Flash-0731 fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency.text-generation0 likes22 downloads22d agoHugging Face15timonziegenbein /fluency-agumented-editstext100K<n<1M0 likes21 downloads11mo agoHugging Face16mkd-minju /keural-v2-fluency-v2 Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning Introduction | Dataset Composition | Methodology | License | Limitations Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.texttext-generation100K<n<1M0 likes20 downloads1mo agoHugging Face17timonziegenbein /fluency-agumented-binarytext100K<n<1M0 likes18 downloads1y agoHugging Face18timonziegenbein /fluency-pairstext10K<n<100K0 likes18 downloads11mo agoHugging Face19toroe /Soofi-German-Fluency-DPO Dataset Card — German Fluency Preference Dataset Overview This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus. All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted. Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-German-Fluency-DPO.texttext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face20timonziegenbein /fluency-editstext10K<n<100K0 likes17 downloads11mo agoHugging Face21speech31 /L2EnglishScoring_speechocean762_fluencyaudion<1K0 likes16 downloads2y agoHugging Face22timonziegenbein /fluency-pairs-gec-only-single-edittext1K<n<10K0 likes15 downloads11mo agoHugging Face23isabelarvelo /fluencybank-3-second-clips Dataset Card for "fluencybank-3-second-clips" More Information needed audio1K<n<10K0 likes14 downloads2y agoHugging Face24tcapelle /completion-coherence-fluencytabular10K<n<100K1 likes13 downloads2y agoHugging Face25isabelarvelo /fluencybank-4-second-clips Dataset Card for "fluencybank-4-second-clips" More Information needed audio1K<n<10K0 likes13 downloads2y agoHugging Face26HosseinRanjbar /FluencyBankaudio0 likes12 downloads10mo agoHugging Face27DynamicSuperb /L2EnglishScoring_speechocean762_fluencyaudion<1K0 likes11 downloads2y agoHugging Face28tcapelle /fluency-datasettext10K<n<100K0 likes11 downloads2y agoHugging Face29timonziegenbein /fluency-edits-gec-only-extendedtext10K<n<100K0 likes11 downloads11mo agoHugging Face30tcapelle /coedit-fluencytext10K<n<100K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.