datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
title-generation-10000x
Title Generation Dataset
Dataset contains 10k prompts and messages collected from over 40 sources with corresponding titles generated by Muse Spark 1.3. Includes 10 languages in english and multilingual splits. Distribution:
Language
Samples
Percentage
English
7000
69.63%
Chinese
417
4.15%
Korean
412
4.10%
Japanese
409
4.07%
Arabic
388
3.86%
French
312
3.10%
Portuguese
294
2.92%
Spanish
292
2.90%
German
266
2.66%
Italian
262
2.61%
task619_ohsumed_abstract_title_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task619_ohsumed_abstract_title_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task619_ohsumed_abstract_title_generation.task219_rocstories_title_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task219_rocstories_title_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task219_rocstories_title_answer_generation.amazon_reviews_multi_fr_prompt_title_generation_from_a_review
amazon_reviews_multi_fr_prompt_title_generation_from_a_review
Summary
amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review.myawady-news-title-generation-dataset
Myawady News Title Generation Dataset 🇲🇲
This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️ Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.orange_sum_fr_prompt_text_generation_from_title_of_an_article
orange_sum_fr_prompt_text_generation_from_title_of_an_article
Summary
orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review.french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 347,688 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset french_book_reviews.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review.orange_sum_fr_prompt_title_generation_from_an_article
orange_sum_fr_prompt_title_generation_from_an_article
Summary
orange_sum_fr_prompt_title_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 639,521 rows that can be used for a title generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_title_generation_from_an_article.Title_Generationads_title_generationDataset to train model
ARGEN_title_generation
Dataset Card for "ARGEN_title_generation"
More Information needed
task1358_xlsum_title_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1358_xlsum_title_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1358_xlsum_title_generation.gulcno_llama2_title_generationtitle-generationtitle_generation
Dataset Card for "title_generation"
More Information needed
chinese_title_generation_gpt_oss_20b
該數據集主要用於訓練模型生成標題
(該數據提取 Mxode/Chinese-Instruct 其中的 5000 條,以及使用 gpt-oss-20b 進行標題生成 (即 response 欄位)。
title_generationflan_source_duorc_ParaphraseRC_title_generation_57flan_source_duorc_SelfRC_title_generation_79llama2_title_generationflan_source_task1161_coda19_title_generation_327flan_source_task1659_title_generation_349flan_combined_task1161_coda19_title_generationflan_combined_task1356_xlsum_title_generationflan_combined_task1358_xlsum_title_generationflan_combined_task418_persent_title_generationflan_combined_task1586_scifact_title_generationdownstream-title-generation
