kth8/title-generation-10000x
Title Generation Dataset Dataset contains 10k prompts and messages collected from over 40 sources with corresponding titles generated by Muse Spark 1.3. Includes 10 languages in english and multilingual splits. Distribution: Language Samples Percentage English 7000 69.63% Chinese 417 4.15% Korean 412 4.10% Japanese 409 4.07% Arabic 388 3.86% French 312 3.10% Portuguese 294 2.92% Spanish 292 2.90% German 266 2.66% Italian 262 2.61%
078
Title Generation Dataset
Dataset contains 10k prompts and messages collected from over 40 sources with corresponding titles generated by Muse Spark 1.3. Includes 10 languages in english and multilingual splits. Distribution:
