CoolFace
Datasetpublic

koalareads/kr-vocab-synth-20260327-v4

KoalaReads German Vocabulary Trainer - Synthetic Conversations Overview This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning. The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference) levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v4.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes60downloads
3 commits on main
7d5ec416mo ago

v4: Fix Parquet format for dataset viewer compatibility

pawelai
e748e856mo ago

v4: LLM-generated conversations replacing templates, judge-filtered (78%+ pass rate)

pawelai
427c0306mo ago

initial commit

pawelai