CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face02RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face03Muennighoff /flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here. Tasks: {'aeslc_10templates', 'ag_news_subset_10templates', 'anli_r1_10templates', 'anli_r2_10templates', 'anli_r3_10templates', 'arc_challenge_10templates', 'arc_easy_10templates', 'bool_q_10templates', 'cb_10templates', 'cnn_dailymail_10templates', 'cola_10templates', 'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.textother1M<n<10M52 likes15k downloads4y agoHugging Face04justachetan /flat-pack-bench Flat-Pack Bench 🧩 Furniture assembly as a spatio-temporal stress test for large vision-language models. Flat-Pack Bench is a multiple-choice benchmark for evaluating fine-grained spatio-temporal understanding in real furniture assembly videos. Each question asks a model to reason about object parts, contact events, assembly order, final connectivity, or part identity across time. Project page: https://flat-pack-bench.github.io 🎯 Benchmark Tasks The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/justachetan/flat-pack-bench.imagevisual-question-answeringn<1K0 likes9.6k downloads4mo agoHugging Face05flaviagiammarino /vqa-rad Dataset Card for VQA-RAD Dataset Description VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from MedPix, which is a free open-access online database of medical images. The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.imagevisual-question-answering1K<n<10K104 likes7.7k downloads3y agoHugging Face06medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.4k downloads3y agoHugging Face07flaviagiammarino /path-vqa Dataset Card for PathVQA Dataset Description PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.imagevisual-question-answering10K<n<100K75 likes6k downloads3y agoHugging Face08chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.3k downloads3y agoHugging Face090xAIT /sinhala-flantext10M<n<100M3 likes3k downloads2y agoHugging Face10FlagEval /EmbSpatial-Bench Introduction Disclaimer: This dataset is organized and adapted from Phineas476/EmbSpatial-Bench. The original data was image format and has been converted here into a more accessible and easy-to-use format. EmbSpatial-Bench is a benchmark for evaluating embodied spatial understanding of LVLMs. The benchmark is automatically derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective. The constructed benchmark comprises a total of 3,640 QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/EmbSpatial-Bench.image1K<n<10K6 likes2.9k downloads1y agoHugging Face11ChanceFocus /flare-finqa Dataset Card for "flare-finqa" More Information needed text1K<n<10K3 likes2.6k downloads3y agoHugging Face12mjaso /flashmini-data-v1 FlashMini data v4 (card) Deterministic FlashMini training corpus. Canonical documents live in Parquet+ZSTD shards under shards/; each shard carries a manifest with sha256, counts, and distributions; the frozen corpus identity is corpus_fingerprint_sha256. Sources and redistribution: each source carries one of mirror_allowed, recipe_only, gated_recipe_only, review_required, generated_owned (fail-closed; see registry/sources.yaml + source_snapshot.lock.json). Content shards are… See the full description on the dataset page: https://huggingface.co/datasets/mjaso/flashmini-data-v1.tabulartext-generation1M<n<10M0 likes2.2k downloads2d agoHugging Face13ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes2.1k downloads2y agoHugging Face14FlagEval /ERQA Introduction Disclaimer: This dataset is organized and adapted from embodiedreasoning/ERQA. The original data was provided in TFRecord format and has been converted here into a more accessible and easy-to-use format. This evaluation benchmark covers a variety of topics related to spatial reasoning and world knowledge focused on real-world scenarios, particularly in the context of robotics. Please find more details and visualizations in the tech report. Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/ERQA.imagen<1K6 likes2k downloads1y agoHugging Face15iraqigold /nih-chest-xray-14-flatimage100K<n<1M0 likes1.9k downloads3mo agoHugging Face16Flame-Code-VLM /Flame-Waterfall-React Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications. The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.textimage-to-text100K<n<1M2 likes1.8k downloads1y agoHugging Face17flaitenberger /wnut_17 WMT-17 This dataset is a Parquet conversion of the original WNT-17 dataset. Source Original authors: Leon Derczynski License: CC BY 4.0 URL: https://huggingface.co/datasets/leondz/wnut_17 Modifications Converted to parquet format Removed arbitrary code execution annotations_creators: crowdsourced language_creators: found language: en license: cc-by-4.0 multilinguality: monolingual size_categories: 1K<n<10K source_datasets: original task_categories:… See the full description on the dataset page: https://huggingface.co/datasets/flaitenberger/wnut_17.text1K<n<10K1 likes1.5k downloads10mo agoHugging Face18flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.4k downloads4y agoHugging Face19SirNeural /flan_v2 Dataset Card for Flan V2 Dataset Summary This is a processed version of the Flan V2 dataset. I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing. The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream. Setup Instructions Here are the steps I followed to get everything working: Build AESLC and WinoGrande datasets… See the full description on the dataset page: https://huggingface.co/datasets/SirNeural/flan_v2.text100M<n<1B200 likes1.4k downloads4y agoHugging Face20abhika-m /fava-flagged-demo Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.textn<1K0 likes1.4k downloads2y agoHugging Face21FlagEval /Where2Place Introduction Disclaimer: This dataset is organized and adapted from wentaoyuan/RoboPoint. The original data was image format and has been converted here into a more accessible and easy-to-use format. This dataset contains 100 real-world images to evaluate free space reference using spatial relations. The images are collected from various cluttered environments. Each image is labeled with a sentence describing the desired some free space and a mask of the desired region.… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/Where2Place.imagen<1K0 likes1.4k downloads1y agoHugging Face22shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.3k downloads1y agoHugging Face23fla-hub /pg19text10K<n<100K0 likes1.3k downloads2y agoHugging Face24clane9 /NSD-Flat NSD-Flat [GitHub] [🤗 Hugging Face Hub] A Hugging Face dataset of pre-processed brain activity flat maps from the Natural Scenes Dataset, constrained to a visual cortex region of interest and rendered as PNG images. Load the dataset Load the dataset from Hugging Face Hub from datasets import load_dataset dataset = load_dataset("clane9/NSD-Flat", split="train") Building the dataset 1. Download source data Run download_data.sh to download the… See the full description on the dataset page: https://huggingface.co/datasets/clane9/NSD-Flat.imageimage-to-image100K<n<1M9 likes1.3k downloads3y agoHugging Face25openguardrails /tb21-dsv4-flash-0731-dsh Terminal-Bench 2.1 trajectories: DeepSeek-V4-Flash-0731 + dsh sdk-minimal Every trial of this one line, in one place: the 89-task main run, both re-run passes, and the scoring scripts. The trajectories are raw and unedited — each step's reasoning, each tool call, and the verifier's own stdout. This is a re-packaging, not a new measurement. The same files were published before, split across two releases, which made the line look incomplete in both: the first release carried the… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-dsv4-flash-0731-dsh.texttext-generation10K<n<100K0 likes1.1k downloads7d agoHugging Face26flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.1k downloads4y agoHugging Face27kowndinya23 /flan2022 Dataset Card for "flan2022" More Information needed text10M<n<100M3 likes956 downloads3y agoHugging Face28malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes953 downloads26d agoHugging Face29flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes936 downloads5y agoHugging Face30ZixuanKe /flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_3text1K<n<10K0 likes866 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.