datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FLAIR-HUB
FLAIR-HUB : Large-scale Multimodal Dataset for Land Cover and Crop Mapping
FLAIR-HUB builds upon and includes the FLAIR#1 and FLAIR#2 datasets, expanding them into a unified, large-scale, multi-sensor land-cover resource with very-high-resolution
annotations. Spanning over 2,500 km² of diverse French ecoclimates and landscapes, it features 63 billion hand-annotated pixels across 19 land-cover and
23 crop type classes.
The dataset integrates complementary data sources including… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR-HUB.FLAIR-1-2
Dataset Card for FLAIR land-cover semantic segmentation
Context & Data
The hereby FLAIR (#1 and #2) dataset is sampled countrywide and is composed of over 20 billion annotated pixels of very high resolution aerial imagery at 0.2 m spatial resolution, acquired over three years and different months (spatio-temporal domains).
Aerial imagery patches consist of 5 channels (RVB-Near Infrared-Elevation) and have corresponding annotation (with 19 semantic classes or 13 for the… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR-1-2.twinicl-bench
TwinICL
38 tasks, each with 132 underlying examples rendered in eight variants: 40,128 rows in total.
Each row contains only:
task: a readable task name.
variant: the text style or image palette.
input_text: the text input, or null for image examples.
input_image: the image input, or null for text examples.
answer: the expected text answer.
The eight variants are lowercase/comma, lowercase/semicolon, uppercase/comma, uppercase/semicolon, and images in neutral, warm, cool, and… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/twinicl-bench.flaird-raid-pan26flair
Federated Learning Annotated Image Repository (FLAIR): A large labelled image dataset for benchmarking in federated learning
FLAIR was published at NeurIPS 2022 (paper)
(Preferred) Benchmarking FLAIR is available in pfl-research (repo, paper).
The ml-flair repo contains a setup for benchmarking with TensorFlow Federated and notebooks for exploring data.
FLAIR is a large dataset of images that captures a number of characteristics encountered in federated learning (FL) and… See the full description on the dataset page: https://huggingface.co/datasets/apple/flair.b-FLAIR-spot
b-FLAIR-spot: bi-temporal extension of FLAIR in SPOT-6/7 modality
Dataset Description
b-FLAIR-spot is a temporal extension of the FLAIR dataset [1], mirroring b-FLAIR in SPOT-6/7 modality, focused on land cover classification in France. The dataset provides bi-temporal satellite image pairs with single-temporal semantic annotations.
Project page: https://xavibou.github.io/CDviaWTS/
Dataset Summary
Task: Semantic change detection via weak temporal… See the full description on the dataset page: https://huggingface.co/datasets/elliotvincent/b-FLAIR-spot.uniref
Dataset Card for flair-bio/uniref
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of
UniRef100 (UniProt Reference Clusters), a comprehensive, non-redundant-by-design database of
protein sequences derived from UniProtKB and select UniParc records. It has been reprocessed by
the FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with additional per-sequence redundancy reduction
(MMseqs2… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/uniref.autotrain-flair-hipe2022-de-hmbert
NER Fine-Tuning
We use Flair for fine-tuning NER models on
HIPE-2022 datasets from
HIPE-2022 Shared Task.
All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from
Lambda Cloud using Flair:
$ git clone https://github.com/flairNLP/flair.git
$ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c
$ pip3 install -e .
$ cd ..
Clone this repo for fine-tuning NER models:
$ git clone https://github.com/stefan-it/hmTEAMS.git
$ cd hmTEAMS/bench
Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-de-hmbert.constructcie
Dataset Card for ConstructCIE
ConstructCIE is a dataset for extracting causal information from construction accident narratives. Each accident report is annotated with a hierarchy of causal factors.
Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Dataset Details
Dataset Description
The dataset contains 530 English construction accident narratives drawn from OSHA accident investigation summaries published… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/constructcie.dewiki-20230701-flair-corpusoas
Dataset Card for flair-bio/oas
Dataset Summary
This dataset is a cleaned and quality-scored version of the paired heavy/light chain subset of
the Observed Antibody Space (OAS) database (Oxford Protein Informatics Group), a large
repository of Next-Generation Sequencing (NGS) antibody repertoires. It has been reprocessed by
the FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded — one shard per source study/run), with per-sequence… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/oas.autotrain-flair-hipe2022-fr-hmbert
NER Fine-Tuning
We use Flair for fine-tuning NER models on
HIPE-2022 datasets from
HIPE-2022 Shared Task.
All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from
Lambda Cloud using Flair:
$ git clone https://github.com/flairNLP/flair.git
$ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c
$ pip3 install -e .
$ cd ..
Clone this repo for fine-tuning NER models:
$ git clone https://github.com/stefan-it/hmTEAMS.git
$ cd hmTEAMS/bench
Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-fr-hmbert.EarthNets_FLAIR2
Dataset Overview
Aerial Imagery
Dimensions: 512 × 512 x 5
Spatial Resolution: 0.2 m
Channels: 5 (RGB, NIR, Elevation)
Sentinel-2 Imagery
Spatial Resolution: 10-20 m
Spectral Bands: 10
Snow/Cloud Masks: Probability range 0-100
Multiple Time Steps: Format T × 10 × W × H (where T, W, H vary)
Labels (Masks)
Dimensions: 512 × 512
Number of Classes: 13
Classes
Class ID
Class Name
Visualization & hint
0
building
🏠
1
pervious… See the full description on the dataset page: https://huggingface.co/datasets/earthflow/EarthNets_FLAIR2.VideoLLM-BoE
VideoLLM-BoE
VideoLLM-BoE contains four benchmarks for evaluating Bag-of-Events behavior in
video large language models. Source code is available in the
GitHub repository.
Dataset contents
Path
Contents
concat-easy/
450 videos and 2,700 questions
concat-hard/
450 videos and 2,700 questions
injected-ads/data.csv
256 examples with questions and MLVU/AdsQA source pairings
natural-ads/data.csv
150 examples with video links, questions, and ad… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/VideoLLM-BoE.flair-multidomain-parquet
FLAIR multi-domain pack — 256 px
24,475 tiles · 9 French domains · 1.94 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: the headline runs; 256 px gives the sharpest boundaries.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes ← this repo
flair-multidomain-parquet-128
128 px
0.9 GB
yes
Columns
column
type
contents
rgb
JPEG bytes
aerial RGB, 256², quality… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet.b-FLAIR
b-FLAIR: bi-temporal extension of FLAIR
Dataset Description
b-FLAIR is a temporal extension of the FLAIR dataset [1] focused on land cover classification in France. The dataset provides bi-temporal orthoimage pairs with single-temporal semantic annotations.
Project page: https://xavibou.github.io/CDviaWTS/
Dataset Summary
Task: Semantic change detection via weak temporal supervision
Coverage: France
Resolution: 0.2 m/px
Patch Size: 512×512 pixels
Bands: Red… See the full description on the dataset page: https://huggingface.co/datasets/elliotvincent/b-FLAIR.flair-multidomain-parquet-512
FLAIR multi-domain pack — 512 px (0.20 m/px)
24,475 tiles · 9 French domains · 4.86 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
This is FLAIR's native aerial resolution. A 102.4 m tile at 512 px is 0.20 m/px — the sampling the COSIA labels were digitised at. The smaller packs below are downsamples, and results from different sampling must not share a table: changing ground sampling changes the task, not just the cost.
Pack
Patch
Ground… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-512.flair-multidomain-parquet-128
FLAIR multi-domain pack — 128 px
24,475 tiles · 9 French domains · 0.9 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: modest GPUs. Same 9 domains and same labels, a quarter of the pixels per training step - the practical choice on a free Colab T4.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes
flair-multidomain-parquet-128
128 px
0.9 GB
yes ← this repo
Columns… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-128.flair_testqa-dataset-k1000
QA Dataset K1000
This dataset contains question-answering data with 1000 distractor documents for long-context evaluation.
Files
File
Description
nq_k1000.jsonl
Natural Questions with 1000 distractors
popqa_k1000.jsonl
PopQA with 1000 distractors
triviaqa_k1000.jsonl
TriviaQA with 1000 distractors
Data Format
Each line is a JSON object with the following fields:
{
"question": "question text",
"answer": ["answer1"… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.mgnify
Dataset Card for flair-bio/mgnify
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of the MGnify
peptide database (EBI Metagenomics), a large collection of predicted protein sequences derived
from environmental metagenomic and metatranscriptomic assemblies. It has been reprocessed by the
FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with per-sequence redundancy reduction (MMseqs2… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/mgnify.flair_transb-FLAIR-test
b-FLAIR-test: Building Change Detection Evaluation Dataset
Dataset Description
b-FLAIR-test is an evaluation dataset for building change detection, containing 1,730 annotated image pairs with binary building change masks. This dataset is designed for in-domain evaluation of methods trained on b-FLAIR in particular and provides a rigorous benchmark for bi-temporal building change detection in general.
Project page: https://xavibou.github.io/CDviaWTS/
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/elliotvincent/b-FLAIR-test.bfd
Dataset Card for flair-bio/bfd
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of the
Big Fantastic Database (BFD) — a large metagenomic protein sequence database originally
assembled from metaclust and used as an MSA source in AlphaFold. It has been reprocessed by the
FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with per-sequence redundancy-reduction (MMseqs2
cascaded… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/bfd.Flair-RSGenFLAIR_1_osm_clip
Dataset Card for "FLAIR_OSM_CLIP"
Dataset for the Seg2Sat model: https://github.com/RubenGres/Seg2Sat
Derived from FLAIR#1 train split.
This dataset incudes the following features:
image: FLAIR#1 .tif files RBG bands converted into a more managable jpg format
segmentation: FLAIR#1 segmentation converted to JPG using the LUT from the documentation
metadata: OSM metadata for the centroid of the image
clip_label: CLIP ViT-H description
class_rep: ratio of appearance of each class in… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR_1_osm_clip.mastermind_24_mcq_randomMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are randomly chosen.
mastermind_35_mcq_closeMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are close to the true solution (by replacing only a single color with a wrong one).
mastermind_46_mcq_closeMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are close to the true solution (by replacing only a single color with a wrong one).
mastermind_46_mcq_randomMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are randomly chosen.
