datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fluency-ecdict-offline
Fluency ECDICT Offline Pack
This repository hosts the versioned ECDICT SQLite packages downloaded by the
Fluency app for offline video word lookup and local subtitle focus-word matching.
Current release
Field
Value
Data version
ecdict-bc015ed2-focus13-v2
Schema
2
Entries
770,611
Focus-word lists
13
ZIP size
54,594,964 bytes
SQLite size
131,756,032 bytes
ZIP SHA-256
1e745ea698878772226a7df584129409dd4b27cb7ea26b4259bc770e7533352f
Source… See the full description on the dataset page: https://huggingface.co/datasets/iamzhangship/fluency-ecdict-offline.fluency-trainerfeedback-prize-english-language-learning-fluencymultilingual-fluencytask138_detoxifying-lms_classification_fluency
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task138_detoxifying-lms_classification_fluency
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task138_detoxifying-lms_classification_fluency.fluency-pairs-gec-onlynormistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Citation
@misc{samuel2025fluentalignmentdisfluentjudges,
title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages},
author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov},
year={2025},
eprint={2512.08777},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.sep28k-fluencybank-stutter-datasetAdaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_AnimalsThe contents of this dataset are from the Forager computer program. Its interactive interface can be found at https://forager.research.bowdoin.edu.
Please read the LICENSE file in this data set and check for further information on the Forager GitHub repository (https://github.com/thelexiconlab/forager-web)
in addition to the docs tab of the web interface linked above.
Citation:
Kumar, A.A., Apsel, M., Zhang, L., Xing, N., Jones. M.N. (2023). forager: A Python package and web interface for… See the full description on the dataset page: https://huggingface.co/datasets/natebranda/Adaptive_Fluency_Similarity_Matrices_and_Frequency_Table_for_Animals.speechocean762_fluencyfluency-edits-geminifluency-pairs-gec-onklyfluency-edits-gec-onlykeural-v2-fluency
Keural-v2 Fluency: A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition (Revisions) |
Methodology |
License |
Limitations
Status: Private staging — not yet cleared for public release (pending §4 evaluation and second-party license audit).
1. Introduction
Keural-v2 Fluency is the "Area 1" component of the Keural-v2 Korean SFT training corpus, purpose-built for DeepSeek-V4-Flash-0731 fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency.fluency-agumented-editskeural-v2-fluency-v2
Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition |
Methodology |
License |
Limitations
Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.fluency-agumented-binaryfluency-pairsSoofi-German-Fluency-DPO
Dataset Card — German Fluency Preference Dataset
Overview
This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus.
All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted.
Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-German-Fluency-DPO.fluency-editsL2EnglishScoring_speechocean762_fluencyfluency-pairs-gec-only-single-editfluencybank-3-second-clips
Dataset Card for "fluencybank-3-second-clips"
More Information needed
completion-coherence-fluencyfluencybank-4-second-clips
Dataset Card for "fluencybank-4-second-clips"
More Information needed
FluencyBankL2EnglishScoring_speechocean762_fluencyfluency-datasetfluency-edits-gec-only-extendedcoedit-fluency
