datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task138_detoxifying-lms_classification_fluency
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task138_detoxifying-lms_classification_fluency
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task138_detoxifying-lms_classification_fluency.keural-v2-fluency-v2
Keural-v2 Fluency (v2): A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition |
Methodology |
License |
Limitations
Status: Private staging — full §3 processing pipeline complete (dedup → PII removal → Korean-purity filter → length check → ratio measurement → train/val/test split → schema normalization → target-model encoding). Pending §4 quantitative/qualitative evaluation and second-party license audit before any… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency-v2.keural-v2-fluency
Keural-v2 Fluency: A Curated Korean–English Conversational Corpus for LLM Fine-Tuning
Introduction |
Dataset Composition (Revisions) |
Methodology |
License |
Limitations
Status: Private staging — not yet cleared for public release (pending §4 evaluation and second-party license audit).
1. Introduction
Keural-v2 Fluency is the "Area 1" component of the Keural-v2 Korean SFT training corpus, purpose-built for DeepSeek-V4-Flash-0731 fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-fluency.Soofi-German-Fluency-DPO
Dataset Card — German Fluency Preference Dataset
Overview
This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus.
All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted.
Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-German-Fluency-DPO.fluencyjfleg_fluency
