CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face02Muennighoff /flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here. Tasks: {'aeslc_10templates', 'ag_news_subset_10templates', 'anli_r1_10templates', 'anli_r2_10templates', 'anli_r3_10templates', 'arc_challenge_10templates', 'arc_easy_10templates', 'bool_q_10templates', 'cb_10templates', 'cnn_dailymail_10templates', 'cola_10templates', 'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.textother1M<n<10M52 likes16k downloads4y agoHugging Face03chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.3k downloads3y agoHugging Face040xAIT /sinhala-flantext10M<n<100M3 likes3.1k downloads2y agoHugging Face05SirNeural /flan_v2 Dataset Card for Flan V2 Dataset Summary This is a processed version of the Flan V2 dataset. I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing. The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream. Setup Instructions Here are the steps I followed to get everything working: Build AESLC and WinoGrande datasets… See the full description on the dataset page: https://huggingface.co/datasets/SirNeural/flan_v2.text100M<n<1B200 likes1.7k downloads4y agoHugging Face06kowndinya23 /flan2022 Dataset Card for "flan2022" More Information needed text10M<n<100M3 likes930 downloads3y agoHugging Face07lorahub /flanv2text1K<n<10K2 likes914 downloads3y agoHugging Face08nguyenthanhdo /FLANv2-without-T0text1M<n<10M1 likes655 downloads2y agoHugging Face09crumb /flan-t5-large-embed-refinedwebAll of the data together is around 81.3GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 1024) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K0 likes649 downloads3y agoHugging Face10coref-data /flan2021_coreference_raw Flan 2021 Coreference Tasks Project: https://github.com/google-research/FLAN/tree/main/flan/v2 Data source: DataProvenanceInitiative/flan2021_submix_original Details This dataset contains all coreference examples that were included in the Flan 2022 collection which were orignally included in Flan 2021. The data is copied from the preprocessed Flan2021 dataset at DataProvenanceInitiative/flan2021_submix_original. COREFERENCE_TASK_NAMES = {… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/flan2021_coreference_raw.text100K<n<1M0 likes612 downloads3y agoHugging Face11sordonia /flan-10k-flat Dataset Card for "flan-10k-flat" More Information needed text10M<n<100M0 likes596 downloads3y agoHugging Face12BEE-spoke-data /flan-v2-hf flan-v2: hf datasets format https://hf.co/datasets/philschmid/flanv2 directly loaded into hf datasets parquet files for easier streaming, etc. [!NOTE] There are two other configs besides the default, en-targets and targets-3w en-targets: the targets column filtered with fasttext-langdetect for en and score >= 0.65 targets-3w: the targets column filtered for 3 or more words text100M<n<1B4 likes539 downloads9mo agoHugging Face13chiayewken /flan-cottext10K<n<100K2 likes464 downloads3y agoHugging Face14r-three /Phatgoose_flanv2_offlinetext1M<n<10M0 likes427 downloads2y agoHugging Face15Lauler /flan-swedishtext1M<n<10M0 likes397 downloads2y agoHugging Face16philschmid /flanv2 Fork of SirNeural/flan_v2 just in case it gets deleted. Dataset Card for Flan V2 Dataset Summary This is a processed version of the Flan V2 dataset. I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing. The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream. This current version I've processed is missing a few… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/flanv2.text1M<n<10M31 likes392 downloads4y agoHugging Face17BEE-spoke-data /FLAN-compressed-plusplus flan-compressed ++ FLAN-compressed with additional tasks added, mostly programming-related. As per Google's flan-v2 README, they seemingly excluded most (all?) programming/code tasks from the data they published. README WIP text100M<n<1B1 likes369 downloads9mo agoHugging Face18lehduong /flan-dolminotext10M<n<100M0 likes335 downloads1y agoHugging Face19aslawliet /flan2021-full Task Name FLAN-2021 -> 70 { "ag_news_subset": 108497, "ai2_arc/ARC-Challenge": 829, "ai2_arc/ARC-Easy": 1927, "aeslc": 13187, "anli/r1": 15361, "anli/r2": 41133, "anli/r3": 91048, "bool_q": 8343, "cnn_dailymail": 259607, "coqa": 6456, "cosmos_qa": 22996, "definite_pronoun_resolution": 1079, "drop": 70045, "fix_punct": 25690, "gem/common_gen": 60936, "gem/dart": 56724, "gem/e2e_nlg": 30337, "gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.texttext-generation10M<n<100M2 likes329 downloads2y agoHugging Face20crumb /flan-t5-base-embed-refinedwebAll of the data together is around 61GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 768) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K1 likes321 downloads3y agoHugging Face21ChamathEka /mini-sinhala-flantabular100K<n<1M6 likes321 downloads2y agoHugging Face22crumb /flan-ul2-tinystoriesAround a quarter of a million examples generated from Flan-UL2 (20b) with the prompt "Write a short story using the vocabulary of a first-grader." to be used in an experimental curriculum learning setting. I had to checkpoint every 1024 examples to mitigate the program slowing down due to memory usage. This was run in bf16 on an RTXA6000 with the following settings: top_k = random between (40, 128) temperature = random between (0.6, 0.95) max_length = 128 batch_size = 32 I wanted a less… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-ul2-tinystories.text100K<n<1M2 likes317 downloads3y agoHugging Face23ai2-adapt-dev /flan_v2_convertedThis is a converted version of the Flan dataset into Tulu SFT training format. The conversion script can be found in our open-instruct repo. The conversion took the following parameters: apply_keyword_filters: True apply_empty_message_filters: True push_to_hub: True hf_entity: ai2-adapt-dev converted_dataset_name: flan_v2_converted local_save_dir: ./data/sft/flan The original FLAN dataset needs extensive efforts to be regenerated, so we are using a reproduced version by the OpenOrca… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/flan_v2_converted.text10K<n<100K3 likes317 downloads2y agoHugging Face24benjamin /flanv2_subsampletext10M<n<100M0 likes313 downloads2y agoHugging Face25crumb /flan-ul2-tinystories-complexAround a quarter of a million examples generated from Flan-UL2 (20b) with the prompt "Write a complex short story using the vocabulary of a third-grader." to be used in an experimental curriculum learning setting. I had to checkpoint every 1024 examples to mitigate the program slowing down due to memory usage. This was run in bf16 on an RTXA6000 with the following settings: top_k = random between (40, 128) temperature = random between (0.6, 0.95) max_length = 128 batch_size = 32 I wanted a… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-ul2-tinystories-complex.text100K<n<1M4 likes309 downloads3y agoHugging Face26vikhyatk /dolmino-mix-1124-flantext10M<n<100M1 likes294 downloads2y agoHugging Face27pszemraj /flan-subsets-deduped flan subsets: deduped [!IMPORTANT] see config all for the aggregated & deduped dataset all configs/subsets have columns inputs and targets deduped on inputs filtered for lang en if contents more than 5 chars. filter out any row with less than 1 char for either column clean-text applied to both columns dedup command deduped with: python -m text_dedup.minhash \ --path $ds_name \ --name $dataset_config \ --split $data_split \ --cache_dir "./cache" \… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/flan-subsets-deduped.text10M<n<100M3 likes284 downloads9mo agoHugging Face28Frejams /Icelandic-Flan Icelandic FLAN Icelandic instruction-following data, built by pairing licensed, human-written Icelandic texts with deterministic instruction templates. Status 16 sources · 46 tasks · 602,057 rows · 45.6M response characters. Source Register Licence Rows Response chars Share umbodsmadur administrative law — Ombudsman art-9 3,914 9,265,216 20.3% igc_news journalism CC BY 4.0 27,711 8,984,257 19.7% rafbokavefur literary — diacritic restoration over… See the full description on the dataset page: https://huggingface.co/datasets/Frejams/Icelandic-Flan.texttext-generation1M<n<10M0 likes265 downloads27d agoHugging Face29crumb /flan-t5-small-embed-refinedwebAll of the data together is around 41GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-small. Structure: { "encoding": List, shaped (512, 512) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } just a tip, you cannot load this with the RAM in the free ver of google colab, not… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-t5-small-embed-refinedweb.textfeature-extraction100K<n<1M0 likes235 downloads3y agoHugging Face30pingzhili /vision-flan_191-task_1k Dataset Card for "vision-flan_191-task_1k" More Information needed image100K<n<1M0 likes229 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.