luth
Datasets
All datasets matching “luth”luth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/luth-sft.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.Luth-2-Post-Training-RL
Luth-2-Post-Training-RL
Luth-2-Post-Training-RL is the French RL prompt collection used to post-train Luth-2-0.8B and Luth-2-2B.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition
Config
Rows
Verifier fields
math
20,000
prompt, solution
math_hard
17,888
prompt, solution
code
46,661
prompt, unit_tests… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-RL.LongSafety-17kLongSafetyBenchluth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/sammybow/luth-sft.
