datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/luth-sft.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.Luth-2-Post-Training-RL
Luth-2-Post-Training-RL
Luth-2-Post-Training-RL is the French RL prompt collection used to post-train Luth-2-0.8B and Luth-2-2B.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition
Config
Rows
Verifier fields
math
20,000
prompt, solution
math_hard
17,888
prompt, solution
code
46,661
prompt, unit_tests… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-RL.luth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/sammybow/luth-sft.luther-bibel-1912
Luther-Bibel (1912) – Deutsche Bibel Public Domain
Beschreibung
Die Lutherbibel 1912 ist die revidierte Fassung von Martin Luthers Bibelübersetzung, herausgegeben von der Deutschen Evangelischen Kirchenkonferenz. Sie ist die maßgebliche deutschsprachige Reformationsbibel und prägte die deutsche Sprache und Theologie entscheidend. Diese Datei enthält alle 66 Bücher in neuer Rechtschreibung, vers-aligniert mit OSIS-Vers-IDs.
Dataset Details
Feld… See the full description on the dataset page: https://huggingface.co/datasets/Geliebter/luther-bibel-1912.
