koalareads/kr-vocab-synth-20260327-v4
KoalaReads German Vocabulary Trainer - Synthetic Conversations Overview This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning. The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference) levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v4.
v4: Fix Parquet format for dataset viewer compatibility
v4: LLM-generated conversations replacing templates, judge-filtered (78%+ pass rate)
initial commit
