datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt4all-j-prompt-generations
Dataset Card for [GPT4All-J Prompt Generations]
Dataset Description
Dataset used to train GPT4All-J and GPT4All-J-LoRA
We release several versions of datasets
v1.0: The original dataset we used to finetune GPT-J on
v1.1-breezy: A filtered dataset where we removed all instances of AI language model
v1.2-jazzy: A filtered dataset where we also removed instances like I'm sorry, I can't answer... and AI language model
v1.3-groovy: The v1.2 dataset with ShareGPT and Dolly… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/gpt4all-j-prompt-generations.gpt4all-j-prompt-generations-pt
Dataset Card for "gpt4all-j-prompt-generations-pt"
Dataset Description
Copy translated into Portuguese of the dataset gpt4all_prompt_generations using the googletrans library.
Translate
translate_dataset.ipynb
Usage
dataset_usage.ipynb
gpt4all_prompt_generations
Dataset Card for [GPT4All Prompt Generations]
Dataset Description
Dataset used to train GPT4All
Homepage:
Repository: gpt4all
Paper: Technical Report
Atlas Map: Map of Cleaned Data
class-to-video-prompt-generationamazon_reviews_multi_fr_prompt_title_generation_from_a_review
amazon_reviews_multi_fr_prompt_title_generation_from_a_review
Summary
amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_binary_text_generation_from_title_of_a_review.newsquadfr_fr_prompt_question_generation_with_answer
newsquadfr_fr_prompt_question_generation_with_answer
Summary
newsquadfr_fr_prompt_question_generation_with_answer is a subset of the Dataset of French Prompts (DFP).It contains 92,620 rows that can be used for a question-generation (with answer) task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_question_generation_with_answer.orange_sum_fr_prompt_text_generation_from_title_of_an_article
orange_sum_fr_prompt_text_generation_from_title_of_an_article
Summary
orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.gpt4all_prompt_generations_with_p3GPT4All extended training set. The original model was trained on https://huggingface.co/datasets/nomic-ai/gpt4all_prompt_generations. We filtered out P3 for our final training. See detail here for why: https://s3.amazonaws.com/static.nomic.ai/gpt4all/2023_GPT4All_Technical_Report.pdf
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review
Summary
amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 7,560,000 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_text_generation_from_title_of_a_review.piaf_fr_prompt_context_generation_with_question
piaf_fr_prompt_context_generation_with_question
Summary
piaf_fr_prompt_context_generation_with_question is a subset of the Dataset of French Prompts (DFP).It contains 442,752 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_context_generation_with_question.piaf_fr_prompt_context_generation_with_answer_and_question
piaf_fr_prompt_context_generation_with_answer_and_question
Summary
piaf_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 442,752 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_context_generation_with_answer_and_question.newsquadfr_fr_prompt_context_generation_with_question
newsquadfr_fr_prompt_context_generation_with_question
Summary
newsquadfr_fr_prompt_context_generation_with_question is a subset of the Dataset of French Prompts (DFP).It contains 101,040 rows that can be used for a context-generation (with question) task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_context_generation_with_question.squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
Summary
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 1,271,928 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question.piaf_fr_prompt_question_generation_with_answer
piaf_fr_prompt_question_generation_with_answer
Summary
piaf_fr_prompt_question_generation_with_answer is a subset of the Dataset of French Prompts (DFP).It contains 387,408 rows that can be used for a question-generation (with answer) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and target columns… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_question_generation_with_answer.UltraChat_50k_Prompt_800k_zephyr_sft_generationsnewsquadfr_fr_prompt_question_generation_with_answer_and_context
newsquadfr_fr_prompt_question_generation_with_answer_and_context
Summary
newsquadfr_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 88,410 rows that can be used for a question generation (with answer and context) task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_question_generation_with_answer_and_context.french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 347,688 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset french_book_reviews.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review.piaf_fr_prompt_question_generation_with_context
piaf_fr_prompt_question_generation_with_context
Summary
piaf_fr_prompt_question_generation_with_context is a subset of the Dataset of French Prompts (DFP).It contains 442,752 rows that can be used for a question-generation (with context) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and target columns… See the full description on the dataset page: https://huggingface.co/datasets/JudSacr/piaf_fr_prompt_question_generation_with_context.Text_2_Graph_Generation_no_prompt-largeOne shot version.
orange_sum_fr_prompt_text_generation_from_an_article
orange_sum_fr_prompt_text_generation_from_an_article
Summary
orange_sum_fr_prompt_text_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 539,400 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_an_article.squad_v2_french_translated_fr_prompt_question_generation_with_context
squad_v2_french_translated_fr_prompt_question_generation_with_context
Summary
squad_v2_french_translated_fr_prompt_question_generation_with_context is a subset of the Dataset of French Prompts (DFP).It contains 3,795,312 rows that can be used for a question-generation (with context) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_question_generation_with_context.paws-x_fr_prompt_paraphrase_generation
paws-x_fr_prompt_paraphrase_generation
Summary
paws-x_fr_prompt_paraphrase_generation is a subset of the Dataset of French Prompts (DFP).It contains 562,728 rows that can be used for a paraphrase generation task.The original data (without prompts) comes from the dataset paws-x by Yang et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/paws-x_fr_prompt_paraphrase_generation.newsquadfr_fr_prompt_context_generation_with_answer_and_question
newsquadfr_fr_prompt_context_generation_with_answer_and_question
Summary
newsquadfr_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 101,040 rows that can be used for a context-generation (with answer)task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_context_generation_with_answer_and_question.orange_sum_fr_prompt_title_generation_from_an_article
orange_sum_fr_prompt_title_generation_from_an_article
Summary
orange_sum_fr_prompt_title_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 639,521 rows that can be used for a title generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_title_generation_from_an_article.squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context
squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context
Summary
squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 1,112,937 rows that can be used for a question-generation (with answer and context) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context.piaf_fr_prompt_question_generation_with_answer_and_context
piaf_fr_prompt_question_generation_with_answer_and_context
Summary
piaf_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 387,408 rows that can be used for a question-generation (with answer and context) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_question_generation_with_answer_and_context.squad_v2_french_translated_fr_prompt_context_generation_with_answer
squad_v2_french_translated_fr_prompt_context_generation_with_answer
Summary
squad_v2_french_translated_fr_prompt_context_generation_with_answer is a subset of the Dataset of French Prompts (DFP).It contains 1,271,928 rows that can be used for a context-generation (with answer) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_context_generation_with_answer.newsquadfr_fr_prompt_context_generation_with_answer
newsquadfr_fr_prompt_context_generation_with_answer
Summary
newsquadfr_fr_prompt_context_generation_with_answer is a subset of the Dataset of French Prompts (DFP).It contains 101,040 rows that can be used for a context-generation (with answer)task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_context_generation_with_answer.llama3_sft_first_corr_prompt_generation4
