datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
title-generation-10000x
Title Generation Dataset
Dataset contains 10k prompts and messages collected from over 40 sources with corresponding titles generated by Muse Spark 1.3. Includes 10 languages in english and multilingual splits. Distribution:
Language
Samples
Percentage
English
7000
69.63%
Chinese
417
4.15%
Korean
412
4.10%
Japanese
409
4.07%
Arabic
388
3.86%
French
312
3.10%
Portuguese
294
2.92%
Spanish
292
2.90%
German
266
2.66%
Italian
262
2.61%
title-generationchinese_title_generation_gpt_oss_20b
該數據集主要用於訓練模型生成標題
(該數據提取 Mxode/Chinese-Instruct 其中的 5000 條,以及使用 gpt-oss-20b 進行標題生成 (即 response 欄位)。
downstream-title-generation
