datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
luth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/luth-sft.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.Luth-2-Post-Training-RL
Luth-2-Post-Training-RL
Luth-2-Post-Training-RL is the French RL prompt collection used to post-train Luth-2-0.8B and Luth-2-2B.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition
Config
Rows
Verifier fields
math
20,000
prompt, solution
math_hard
17,888
prompt, solution
code
46,661
prompt, unit_tests… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-RL.LongSafety-17kLongSafetyBenchluth-sft
Dataset Details
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
Dataset Sources
Scholar
By Kurakura AI: Dataset Link.Built… See the full description on the dataset page: https://huggingface.co/datasets/sammybow/luth-sft.luther-bibel-1912
Luther-Bibel (1912) – Deutsche Bibel Public Domain
Beschreibung
Die Lutherbibel 1912 ist die revidierte Fassung von Martin Luthers Bibelübersetzung, herausgegeben von der Deutschen Evangelischen Kirchenkonferenz. Sie ist die maßgebliche deutschsprachige Reformationsbibel und prägte die deutsche Sprache und Theologie entscheidend. Diese Datei enthält alle 66 Bücher in neuer Rechtschreibung, vers-aligniert mit OSIS-Vers-IDs.
Dataset Details
Feld… See the full description on the dataset page: https://huggingface.co/datasets/Geliebter/luther-bibel-1912.luther_bibel_1545_de
Lutherbibel (1545)
Description
The original Luther Bible translation by Martin Luther (1483-1546), first published in 1534. This 1545 edition represents Luther's final revision of his landmark translation. Luther's translation of the Bible into German from the original Hebrew and Greek was a watershed moment of the Reformation and had a profound influence on the development of the modern German language. It includes the Protestant canon (66 books).… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/luther_bibel_1545_de.luther_bibel_de
Lutherbibel (1912)
Description
The Luther Bible (Lutherbibel) is a German translation of the Bible by Martin Luther (1483-1546), first published in 1534. This 1912 edition represents the final revision of Luther's translation before the major revisions of the 20th century. Luther's translation of the Bible into German from the original Hebrew and Greek was a landmark of the Reformation and had a profound influence on the development of the German language. The… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/luther_bibel_de.fine-llama-2-testsample-construction-datasetv1_modelFood_Recognitiontourism-package-predictionData files for the tourism package prediction project.
