CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face02Muennighoff /flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here. Tasks: {'aeslc_10templates', 'ag_news_subset_10templates', 'anli_r1_10templates', 'anli_r2_10templates', 'anli_r3_10templates', 'arc_challenge_10templates', 'arc_easy_10templates', 'bool_q_10templates', 'cb_10templates', 'cnn_dailymail_10templates', 'cola_10templates', 'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.textother1M<n<10M52 likes16k downloads4y agoHugging Face03orionweller /tulu_flan_mds_incremental-tokens0 likes4.4k downloads2y agoHugging Face04Vision-Flan /vision-flan_191-task_1k 🚀 Vision-Flan Dataset vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task. It is constructed for visual instruction tuning and for building large-scale vision-language models. Paper or blog for more information: https://github.com/VT-NLP/MultiInstruct/ https://vision-flan.github.io/ Paper coming soon 😊 Citation Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.imagevisual-question-answering100K<n<1M22 likes3.6k downloads3y agoHugging Face05orionweller /tulu_flan_mds_incremental0 likes3.4k downloads2y agoHugging Face06chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.3k downloads3y agoHugging Face070xAIT /sinhala-flantext10M<n<100M3 likes3.1k downloads2y agoHugging Face08SirNeural /flan_v2 Dataset Card for Flan V2 Dataset Summary This is a processed version of the Flan V2 dataset. I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing. The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream. Setup Instructions Here are the steps I followed to get everything working: Build AESLC and WinoGrande datasets… See the full description on the dataset page: https://huggingface.co/datasets/SirNeural/flan_v2.text100M<n<1B200 likes1.7k downloads4y agoHugging Face09internlm /Agent-FLAN Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models This page holds the dataset proposed in Agent-FLAN, which consists of AgentInstruct, Toolbench, and customized negative agent samples as its source datasets. ✨ Introduction [🤗 HuggingFace] [📃 Paper] [🌐 Project Page] Open-sourced Large Language Models (LLMs) have achieved great success in various NLP tasks, however, they are still far inferior to API-based models when acting as… See the full description on the dataset page: https://huggingface.co/datasets/internlm/Agent-FLAN.106 likes1.4k downloads3y agoHugging Face10HKUSTAudio /Audio-FLAN-Datasetgated Audio-FLAN Dataset (Paper) (the FULL audio files and jsonl files are still updating) An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound. 1. Dataset Structure The Audio-FLAN-Dataset has the following directory structure: Audio-FLAN-Dataset/ ├── audio_files/ │ ├── audio/ │ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/ │ │ └── 179_Audioset_for_Audio_Inpainting/ │ │ └── ... │ ├── music/ │ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.audiotext-to-speech10M<n<100M47 likes1.2k downloads1y agoHugging Face11kowndinya23 /flan2022 Dataset Card for "flan2022" More Information needed text10M<n<100M3 likes930 downloads3y agoHugging Face12lorahub /flanv2text1K<n<10K2 likes914 downloads3y agoHugging Face13Vision-Flan /vision-flan Image generated by https://ideogram.ai/ We introduce Vision-Flan, the largest human-annotated visual instruction tuning dataset that consists of 200+ diverse vision-language tasks derived from 101 open-source computer vision datasets. Each task is equipped with an expert written instruction and carefully designed templates for the inputs and outputs. The dataset encompasses a wide range of tasks such as image captioning, visual question-answering, and visual understanding. Vision-Flan is… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan.image1K<n<10K7 likes696 downloads2y agoHugging Face14nguyenthanhdo /FLANv2-without-T0text1M<n<10M1 likes655 downloads2y agoHugging Face15crumb /flan-t5-large-embed-refinedwebAll of the data together is around 81.3GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 1024) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K0 likes649 downloads3y agoHugging Face16coref-data /flan2021_coreference_raw Flan 2021 Coreference Tasks Project: https://github.com/google-research/FLAN/tree/main/flan/v2 Data source: DataProvenanceInitiative/flan2021_submix_original Details This dataset contains all coreference examples that were included in the Flan 2022 collection which were orignally included in Flan 2021. The data is copied from the preprocessed Flan2021 dataset at DataProvenanceInitiative/flan2021_submix_original. COREFERENCE_TASK_NAMES = {… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/flan2021_coreference_raw.text100K<n<1M0 likes612 downloads3y agoHugging Face17sordonia /flan-10k-flat Dataset Card for "flan-10k-flat" More Information needed text10M<n<100M0 likes596 downloads3y agoHugging Face18BEE-spoke-data /flan-v2-hf flan-v2: hf datasets format https://hf.co/datasets/philschmid/flanv2 directly loaded into hf datasets parquet files for easier streaming, etc. [!NOTE] There are two other configs besides the default, en-targets and targets-3w en-targets: the targets column filtered with fasttext-langdetect for en and score >= 0.65 targets-3w: the targets column filtered for 3 or more words text100M<n<1B4 likes539 downloads9mo agoHugging Face19chiayewken /flan-cottext10K<n<100K2 likes464 downloads3y agoHugging Face20BJyotibrat /ROBIN-ImagesGT-Merged-Flanora-AI-v1 Flanora AI/ROBIN-ImagesGT-Merged-Flanora-AI-v1 ROBIN-ImagesGT-Merged-Flanora-AI-v1 is a curated collection of 622 floor-plan images created by merging floor-plan data from the ROBIN dataset and the CVC-FP / ImagesGT dataset. The dataset is organized into six bedroom-count categories: 0 bedroom, 1 bedroom, 2 bedroom, 3 bedroom, 4 bedroom, and 5 bedroom. The dataset contains the original, unprocessed floor-plan images. No image preprocessing or transformation was applied to the… See the full description on the dataset page: https://huggingface.co/datasets/BJyotibrat/ROBIN-ImagesGT-Merged-Flanora-AI-v1.imageimage-to-imagen<1K1 likes429 downloads1mo agoHugging Face21r-three /Phatgoose_flanv2_offlinetext1M<n<10M0 likes427 downloads2y agoHugging Face22Lauler /flan-swedishtext1M<n<10M0 likes397 downloads2y agoHugging Face23philschmid /flanv2 Fork of SirNeural/flan_v2 just in case it gets deleted. Dataset Card for Flan V2 Dataset Summary This is a processed version of the Flan V2 dataset. I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing. The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream. This current version I've processed is missing a few… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/flanv2.text1M<n<10M31 likes392 downloads4y agoHugging Face24BEE-spoke-data /FLAN-compressed-plusplus flan-compressed ++ FLAN-compressed with additional tasks added, mostly programming-related. As per Google's flan-v2 README, they seemingly excluded most (all?) programming/code tasks from the data they published. README WIP text100M<n<1B1 likes369 downloads9mo agoHugging Face25lehduong /flan-dolminotext10M<n<100M0 likes335 downloads1y agoHugging Face26aslawliet /flan2021-full Task Name FLAN-2021 -> 70 { "ag_news_subset": 108497, "ai2_arc/ARC-Challenge": 829, "ai2_arc/ARC-Easy": 1927, "aeslc": 13187, "anli/r1": 15361, "anli/r2": 41133, "anli/r3": 91048, "bool_q": 8343, "cnn_dailymail": 259607, "coqa": 6456, "cosmos_qa": 22996, "definite_pronoun_resolution": 1079, "drop": 70045, "fix_punct": 25690, "gem/common_gen": 60936, "gem/dart": 56724, "gem/e2e_nlg": 30337, "gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.texttext-generation10M<n<100M2 likes329 downloads2y agoHugging Face27crumb /flan-t5-base-embed-refinedwebAll of the data together is around 61GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 768) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K1 likes321 downloads3y agoHugging Face28ChamathEka /mini-sinhala-flantabular100K<n<1M6 likes321 downloads2y agoHugging Face29crumb /flan-ul2-tinystoriesAround a quarter of a million examples generated from Flan-UL2 (20b) with the prompt "Write a short story using the vocabulary of a first-grader." to be used in an experimental curriculum learning setting. I had to checkpoint every 1024 examples to mitigate the program slowing down due to memory usage. This was run in bf16 on an RTXA6000 with the following settings: top_k = random between (40, 128) temperature = random between (0.6, 0.95) max_length = 128 batch_size = 32 I wanted a less… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-ul2-tinystories.text100K<n<1M2 likes317 downloads3y agoHugging Face30ai2-adapt-dev /flan_v2_convertedThis is a converted version of the Flan dataset into Tulu SFT training format. The conversion script can be found in our open-instruct repo. The conversion took the following parameters: apply_keyword_filters: True apply_empty_message_filters: True push_to_hub: True hf_entity: ai2-adapt-dev converted_dataset_name: flan_v2_converted local_save_dir: ./data/sft/flan The original FLAN dataset needs extensive efforts to be regenerated, so we are using a reproduced version by the OpenOrca… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/flan_v2_converted.text10K<n<100K3 likes317 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.