madmoonslice/German-Instruct-Icarus
German-Instruct-Icarus Dataset Overview The German-Instruct-Icarus dataset is a high-quality, fully copyright-free collection of 10 million German-language examples, carefully curated for instruction tuning, fine-tuning, and pre-training of large language models (LLMs). It emphasizes natural dialogue, conversational patterns, and diverse textual interactions. All data is sourced exclusively from public domain and openly licensed materials, including transcribed… See the full description on the dataset page: https://huggingface.co/datasets/madmoonslice/German-Instruct-Icarus.
This repository belongs to madmoonslice on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
