datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaia_filtered_text_onlyvocab_filtered_dataset_22B
Dataset Card for "vocab_filtered_dataset_22B"
Dataset Summary
This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES)
We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.vocab_filtered_dataset_2.1B
Dataset Card for "vocab_filtered_dataset_2.1B"
Dataset Summary
This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES)
We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_2.1B.GigaVerbo-Text-Filter
GigaVerbo Text-Filter
Dataset Summary
GigaVerbo Text-Filter is a dataset with 110,000 randomly selected samples from 9 subsets of GigaVerbo (i.e., specifically those that were not synthetic). This dataset was used to train the text-quality filters described in "Tucano: Advancing Neural Text Generation for Portuguese". To create the text embeddings, we used sentence-transformers/LaBSE. All scores were generated by GPT-4o.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/TucanoBR/GigaVerbo-Text-Filter.instruction-data-text-only-multiturn-filtered-for-tokenizearabic-quran-filtered-textinstruction-data-text-only-multiturn-filtered-for-tokenize-cleansynthetic_text_to_sql_filterVQAv2_sample_validation_text_davinci_002_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_2
Dataset Card for "VQAv2_sample_validation_text_davinci_002_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_2"
More Information needed
OK-VQA_test_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_100
Dataset Card for "OK-VQA_test_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_100"
More Information needed
VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_100
Dataset Card for "VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_100"
More Information needed
VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_200
Dataset Card for "VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_200"
More Information needed
VQAv2_sample_validation_text_davinci_002_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_10
Dataset Card for "VQAv2_sample_validation_text_davinci_002_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_10"
More Information needed
VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_10
Dataset Card for "VQAv2_sample_validation_text_davinci_003_mode_T_A_D_PNP_NO_FILTER_C_Q_rices_ns_10"
More Information needed
Magicoder-Evol-Instruct-110K-Filtered_0.35-textfiltered_Dogs_image_text_pairtasariv_splits_transcibed_filtered-tags-and-text-generatedtext_restyle_filtered
