datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flan-t5-boosting-mmlu_cotflan-t5-boosting-bbh_cotflan-t5-boosting-bbh_directflan-t5-boosting-tydiqa_directflan-t5-boosting-bbh_zstext_to_sql_FLANflan-t5-boosting-tydiqa_zsflan-t5-boosting-mgsm_cotflanv2_cot_dedepulicated
FLAN v2 Cot Deduplicated Dataset
Data Preprocessing
Remove instructions with less than 100 tokens in 'targets'.
Dedepulicate Dataset using cosine similarity with a threshold of 0.95.
Code
Github repo : https://github.com/AJlearner46/Deduplicate-flanv2-finetune-LLaMa3-
Acknowledgments
The original dataset is provided by SirNeural/flan_v2.
Tokenizer used: bert-base-uncased from Hugging Face.
pronto-qa-flanT5custom_flan_T5_Datasetflan-t5-boosting-mmlu_directflan-t5-boosting-mgsm_zsFLAN-IntentFlanv2_cot_dataset_4_APKflanv2-cotemail-priority-flant5flan-t5-boosting-benchmarksFLAN_V2_dataset_APKFLANV2_newdataflan_cot
Cleaned FLAN v2 Dataset 8K
Dataset Name: Cleaned FLAN v2 DatasetSource: SirNeural/flan_v2Files Included:
cot_fs_noopt_train.jsonl
cot_fs_opt_train.jsonl
cot_zs_noopt_train.jsonl
cot_zs_opt_train.jsonl
Data Processing and Cleaning
The dataset was processed and cleaned using the following steps:
Merging the JSONL Files:
The original dataset comprises four separate JSONL files:
cot_fs_noopt_train.jsonl
cot_fs_opt_train.jsonl
cot_zs_noopt_train.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/KameshRsk/flan_cot.flanv2modifiedflan-2024Flanv2_cot_dataset_APKflan_v2_datasetThis Dataset is used for LLaMA-2 model to train
oath-frames-flan-datasetsFlan-V2-Submix-2024FlanV2-2024
