title-generation
title-generation-10000x
Title Generation Dataset
Dataset contains 10k prompts and messages collected from over 40 sources with corresponding titles generated by Muse Spark 1.3. Includes 10 languages in english and multilingual splits. Distribution:
Language
Samples
Percentage
English
7000
69.63%
Chinese
417
4.15%
Korean
412
4.10%
Japanese
409
4.07%
Arabic
388
3.86%
French
312
3.10%
Portuguese
294
2.92%
Spanish
292
2.90%
German
266
2.66%
Italian
262
2.61%
task619_ohsumed_abstract_title_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task619_ohsumed_abstract_title_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task619_ohsumed_abstract_title_generation.task219_rocstories_title_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task219_rocstories_title_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task219_rocstories_title_answer_generation.amazon_reviews_multi_fr_prompt_title_generation_from_a_review
amazon_reviews_multi_fr_prompt_title_generation_from_a_review
Summary
amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review.myawady-news-title-generation-dataset
Myawady News Title Generation Dataset 🇲🇲
This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️ Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.
