datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FLAN🍮 The WHOLE FLAN Collection! 🍮
Overview
This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets.
Generated using the official seqio templating from the Google FLAN Collection GitHub repo.
The data is subject to all the same licensing of the component datasets.
To keep up with our continued work on OpenOrca and other exciting research, find our Discord here:
https://AlignmentLab.ai
Motivation
This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here.
Tasks:
{'aeslc_10templates',
'ag_news_subset_10templates',
'anli_r1_10templates',
'anli_r2_10templates',
'anli_r3_10templates',
'arc_challenge_10templates',
'arc_easy_10templates',
'bool_q_10templates',
'cb_10templates',
'cnn_dailymail_10templates',
'cola_10templates',
'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.flat-pack-bench
Flat-Pack Bench 🧩
Furniture assembly as a spatio-temporal stress test for large vision-language models.
Flat-Pack Bench is a multiple-choice benchmark for evaluating fine-grained
spatio-temporal understanding in real furniture assembly videos. Each question
asks a model to reason about object parts, contact events, assembly order, final
connectivity, or part identity across time.
Project page: https://flat-pack-bench.github.io
🎯 Benchmark Tasks
The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/justachetan/flat-pack-bench.vqa-rad
Dataset Card for VQA-RAD
Dataset Description
VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from MedPix, which is a free open-access online database of medical images.
The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.medical_meadow_medical_flashcards
Dataset Card for Medical Flashcards
Dataset Summary
Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master
in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge,
and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the
entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.path-vqa
Dataset Card for PathVQA
Dataset Description
PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a
publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.flan-v2
Dataset Card for "flan-v2"
More Information needed
sinhala-flanEmbSpatial-Bench
Introduction
Disclaimer: This dataset is organized and adapted from Phineas476/EmbSpatial-Bench. The original data was image format and has been converted here into a more accessible and easy-to-use format.
EmbSpatial-Bench is a benchmark for evaluating embodied spatial understanding of LVLMs. The benchmark is automatically derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective. The constructed benchmark comprises a total of 3,640 QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/EmbSpatial-Bench.flare-finqa
Dataset Card for "flare-finqa"
More Information needed
flashmini-data-v1
FlashMini data v4 (card)
Deterministic FlashMini training corpus. Canonical documents live in
Parquet+ZSTD shards under shards/; each shard carries a manifest with
sha256, counts, and distributions; the frozen corpus identity is
corpus_fingerprint_sha256.
Sources and redistribution: each source carries one of mirror_allowed,
recipe_only, gated_recipe_only, review_required, generated_owned
(fail-closed; see registry/sources.yaml + source_snapshot.lock.json).
Content shards are… See the full description on the dataset page: https://huggingface.co/datasets/mjaso/flashmini-data-v1.ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitERQA
Introduction
Disclaimer: This dataset is organized and adapted from embodiedreasoning/ERQA. The original data was provided in TFRecord format and has been converted here into a more accessible and easy-to-use format.
This evaluation benchmark covers a variety of topics related to spatial reasoning and world knowledge focused on real-world scenarios, particularly in the context of robotics. Please find more details and visualizations in the tech report.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/ERQA.nih-chest-xray-14-flatFlame-Waterfall-React
Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation
Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications.
The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.wnut_17
WMT-17
This dataset is a Parquet conversion of the original WNT-17 dataset.
Source
Original authors: Leon Derczynski
License: CC BY 4.0
URL: https://huggingface.co/datasets/leondz/wnut_17
Modifications
Converted to parquet format
Removed arbitrary code execution
annotations_creators:
crowdsourced
language_creators:
found
language:
en
license:
cc-by-4.0
multilinguality:
monolingual
size_categories:
1K<n<10K
source_datasets:
original
task_categories:… See the full description on the dataset page: https://huggingface.co/datasets/flaitenberger/wnut_17.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.flan_v2
Dataset Card for Flan V2
Dataset Summary
This is a processed version of the Flan V2 dataset.
I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing.
The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream.
Setup Instructions
Here are the steps I followed to get everything working:
Build AESLC and WinoGrande datasets… See the full description on the dataset page: https://huggingface.co/datasets/SirNeural/flan_v2.fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.Where2Place
Introduction
Disclaimer: This dataset is organized and adapted from wentaoyuan/RoboPoint. The original data was image format and has been converted here into a more accessible and easy-to-use format.
This dataset contains 100 real-world images to evaluate free space reference using spatial relations. The images are collected from various cluttered environments. Each image is labeled with a sentence describing the desired some free space and a mask of the desired region.… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/Where2Place.gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.pg19NSD-Flat
NSD-Flat
[GitHub] [🤗 Hugging Face Hub]
A Hugging Face dataset of pre-processed brain activity flat maps from the Natural Scenes Dataset, constrained to a visual cortex region of interest and rendered as PNG images.
Load the dataset
Load the dataset from Hugging Face Hub
from datasets import load_dataset
dataset = load_dataset("clane9/NSD-Flat", split="train")
Building the dataset
1. Download source data
Run download_data.sh to download the… See the full description on the dataset page: https://huggingface.co/datasets/clane9/NSD-Flat.tb21-dsv4-flash-0731-dsh
Terminal-Bench 2.1 trajectories: DeepSeek-V4-Flash-0731 + dsh sdk-minimal
Every trial of this one line, in one place: the 89-task main run, both re-run passes, and
the scoring scripts. The trajectories are raw and unedited — each step's reasoning, each
tool call, and the verifier's own stdout.
This is a re-packaging, not a new measurement. The same files were published before,
split across two releases, which made the line look incomplete in both: the first release
carried the… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-dsv4-flash-0731-dsh.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.flan2022
Dataset Card for "flan2022"
More Information needed
GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_3
