datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twinicl-bench
TwinICL
38 tasks, each with 132 underlying examples rendered in eight variants: 40,128 rows in total.
Each row contains only:
task: a readable task name.
variant: the text style or image palette.
input_text: the text input, or null for image examples.
input_image: the image input, or null for text examples.
answer: the expected text answer.
The eight variants are lowercase/comma, lowercase/semicolon, uppercase/comma, uppercase/semicolon, and images in neutral, warm, cool, and… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/twinicl-bench.flaird-raid-pan26flair
Federated Learning Annotated Image Repository (FLAIR): A large labelled image dataset for benchmarking in federated learning
FLAIR was published at NeurIPS 2022 (paper)
(Preferred) Benchmarking FLAIR is available in pfl-research (repo, paper).
The ml-flair repo contains a setup for benchmarking with TensorFlow Federated and notebooks for exploring data.
FLAIR is a large dataset of images that captures a number of characteristics encountered in federated learning (FL) and… See the full description on the dataset page: https://huggingface.co/datasets/apple/flair.uniref
Dataset Card for flair-bio/uniref
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of
UniRef100 (UniProt Reference Clusters), a comprehensive, non-redundant-by-design database of
protein sequences derived from UniProtKB and select UniParc records. It has been reprocessed by
the FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with additional per-sequence redundancy reduction
(MMseqs2… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/uniref.constructcie
Dataset Card for ConstructCIE
ConstructCIE is a dataset for extracting causal information from construction accident narratives. Each accident report is annotated with a hierarchy of causal factors.
Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Dataset Details
Dataset Description
The dataset contains 530 English construction accident narratives drawn from OSHA accident investigation summaries published… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/constructcie.dewiki-20230701-flair-corpusoas
Dataset Card for flair-bio/oas
Dataset Summary
This dataset is a cleaned and quality-scored version of the paired heavy/light chain subset of
the Observed Antibody Space (OAS) database (Oxford Protein Informatics Group), a large
repository of Next-Generation Sequencing (NGS) antibody repertoires. It has been reprocessed by
the FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded — one shard per source study/run), with per-sequence… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/oas.flair-multidomain-parquet
FLAIR multi-domain pack — 256 px
24,475 tiles · 9 French domains · 1.94 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: the headline runs; 256 px gives the sharpest boundaries.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes ← this repo
flair-multidomain-parquet-128
128 px
0.9 GB
yes
Columns
column
type
contents
rgb
JPEG bytes
aerial RGB, 256², quality… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet.flair-multidomain-parquet-512
FLAIR multi-domain pack — 512 px (0.20 m/px)
24,475 tiles · 9 French domains · 4.86 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
This is FLAIR's native aerial resolution. A 102.4 m tile at 512 px is 0.20 m/px — the sampling the COSIA labels were digitised at. The smaller packs below are downsamples, and results from different sampling must not share a table: changing ground sampling changes the task, not just the cost.
Pack
Patch
Ground… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-512.flair-multidomain-parquet-128
FLAIR multi-domain pack — 128 px
24,475 tiles · 9 French domains · 0.9 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: modest GPUs. Same 9 domains and same labels, a quarter of the pixels per training step - the practical choice on a free Colab T4.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes
flair-multidomain-parquet-128
128 px
0.9 GB
yes ← this repo
Columns… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-128.flair_testqa-dataset-k1000
QA Dataset K1000
This dataset contains question-answering data with 1000 distractor documents for long-context evaluation.
Files
File
Description
nq_k1000.jsonl
Natural Questions with 1000 distractors
popqa_k1000.jsonl
PopQA with 1000 distractors
triviaqa_k1000.jsonl
TriviaQA with 1000 distractors
Data Format
Each line is a JSON object with the following fields:
{
"question": "question text",
"answer": ["answer1"… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.mgnify
Dataset Card for flair-bio/mgnify
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of the MGnify
peptide database (EBI Metagenomics), a large collection of predicted protein sequences derived
from environmental metagenomic and metatranscriptomic assemblies. It has been reprocessed by the
FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with per-sequence redundancy reduction (MMseqs2… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/mgnify.flair_transbfd
Dataset Card for flair-bio/bfd
Dataset Summary
This dataset is a cleaned, deduplication-clustered, and quality-scored version of the
Big Fantastic Database (BFD) — a large metagenomic protein sequence database originally
assembled from metaclust and used as an MSA source in AlphaFold. It has been reprocessed by the
FLAIR modules/data pipeline into a single
training-ready Parquet dataset (sharded), with per-sequence redundancy-reduction (MMseqs2
cascaded… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/bfd.FLAIR_1_osm_clip
Dataset Card for "FLAIR_OSM_CLIP"
Dataset for the Seg2Sat model: https://github.com/RubenGres/Seg2Sat
Derived from FLAIR#1 train split.
This dataset incudes the following features:
image: FLAIR#1 .tif files RBG bands converted into a more managable jpg format
segmentation: FLAIR#1 segmentation converted to JPG using the LUT from the documentation
metadata: OSM metadata for the centroid of the image
clip_label: CLIP ViT-H description
class_rep: ratio of appearance of each class in… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR_1_osm_clip.mastermind_24_mcq_randomMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are randomly chosen.
mastermind_35_mcq_closeMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are close to the true solution (by replacing only a single color with a wrong one).
mastermind_46_mcq_closeMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are close to the true solution (by replacing only a single color with a wrong one).
mastermind_46_mcq_randomMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are randomly chosen.
mastermind_24_mcq_closeMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are close to the true solution (by replacing only a single color with a wrong one).
mastermind_35_mcq_randomMultiple-choice reasoning benchmark based on the game of Mastermind. Answer options are randomly chosen.
flair2
GeoBench-2 Dataset License Attribution
Dataset Name: m-FLAIR-2Original Dataset Name: FLAIR #2Original Source: https://huggingface.co/datasets/IGNF/FLAIR-1-2
Related Publication(s): https://arxiv.org/abs/2310.13336
Licensing
Annotation License: Open Licence 2.0 (Etalab)
Image License: Copernicus Open Access for Sentinel-2 + IGN aerial imagery under Etalab / public domain as per IGN’s open data policy
Declared By Original Provider: IGN / FLAIR project pages, GitHub… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/flair2.nuner_max_splitsflair-testmastermind_35_promptflair-test-2pilener_entropy_splitsmastermind_46_promptmastermind_24_prompt
