datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DFP
Dataset Card for Dataset of French Prompts (DFP)
This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/DFP.DFP
Dataset Card for Dataset of French Prompts (DFP)
This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/JudSacr/DFP.DFPO-Preft-adelieDFPO-Preft-taiyiDFPO-MFPEAdf_persona_final_v6_undersampleddf_persona_final_v6_unbalanced_v2_qwen7bInstruct_comrag_sep_kedf_persona_final_v6_unbalanced_v2_qwen7bInstruct_comrag_sep_geral_kedf_persona_final_v6_unbalanced_v2_600df_persona_final_v6_unbalanced_v2_qwen7bInstruct_alldf_persona_final_v6_unbalanced_v2_llama3b_alldf_persona_final_v6_unbalanced_v2_qwen7bInstruct_norag_ke
