datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.conceptual-12m-mbart-50-multilingualconceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
food17-flaggedconceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following:
train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each)
val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following:
train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.flan-t5-boosting-mmlu_cotsurya-bench-flare-forecasting
Full-disk Solar Flare Forecasting Dataset
Dataset Summary
This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index.
Two derived binary labels are included for forecasting tasks:
label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.evalsconceptual-12m-multilingual-marian-esafrica-synth-energy-oilgas-gas-flaring-nigeria
Africa Synth Energy Oilgas Gas Flaring Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-flaring-nigeria.spanish-nahuatl-flaggingmultilingual-vqaskin-cancer-flagged-dataset
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AreejAlotaibi12/skin-cancer-flagged-dataset.CIL-Dataset
CIL-Dataset: Colombian Indigenous Languages
Parallel and monolingual corpora for machine translation between Spanish and
four Colombian Indigenous languages: Wayuunaiki, Inga, Kamëntsá and
Nasa Yuwe. Compiled for the thesis Data-Centric Strategies for Low-Resource
Neural Machine Translation in Colombian Indigenous Languages (Universidad de los
Andes, FLAG Lab).
Languages
Language
Column
ISO 639-3
Spanish
esp
spa
Wayuunaiki
way
guc
Inga
ing
inb… See the full description on the dataset page: https://huggingface.co/datasets/Flaglab/CIL-Dataset.agronomy-qa-agriculture
Agronomy QA Pairs — Agriculture Instruction Dataset
A concise, English-language question-and-answer dataset covering practical agriculture:
crop management and planting, soil health and fertility, irrigation, and pest and disease
control, with a smaller amount of livestock content. Curated for instruction fine-tuning
of language models in the agriculture domain.
Rows
1,914 question/answer pairs
Unique questions
1,914 — no duplicate questions
Unique answers
1,870… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/agronomy-qa-agriculture.flare-fomcflat-birthday-cf2c09
flat-birthday-cf2c09
Synthetic products test data: 41 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/frostglade/flat-birthday-cf2c09.market-news-qa
Market News QA — Market-Analysis Instruction Dataset
Concise question-and-answer pairs for market analysis and financial news
interpretation: classifying news by market area, reading sentiment, identifying
who a story matters to, and answering forward-looking questions from earnings calls.
Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category).
Rows
10,011
Distinct answers
8,469 (85%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.flan-t5-boosting-bbh_cotindian-skincare-ingredient-flags
Indian Skincare Ingredient Flags
1,448 cosmetic (INCI) ingredients found in skincare sold in India, with dermatology-referenced
flags: Malassezia (fungal-acne) trigger status, comedogenicity rating, fragrance/EU-allergen,
alcohol type and pregnancy-caution.
Published by CureSkin (Cure and Care Wellness Pvt. Ltd.) — the data layer
behind Indian Skincare, Decoded, medically reviewed
by Dr. Charu Sharma, Head of Dermatology at CureSkin.
Blank cells mean "not assessed", never "safe"… See the full description on the dataset page: https://huggingface.co/datasets/cureskin/indian-skincare-ingredient-flags.Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias.
The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations.
The Structure of the dataset is of the following type:
Base Sentence
Occupation
Steretypical_Gender
Male Sentence
Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.flagged-antibody-catalogs
Flagged antibody catalogs
A one-row-per-product view of the 2026 antibody validation-image audit: which catalog
numbers have at least one problematic validation image on the vendor's own product page.
Source and credit
Derived from Richardson, R. & David, S. (2026), "Problematic images in vendor antibody
verification data", Zenodo, version 260825, published 25 August 2026, DOI
10.5281/zenodo.22090940, licensed CC BY 4.0.
The source audit documents 18,943… See the full description on the dataset page: https://huggingface.co/datasets/wren-ilands/flagged-antibody-catalogs.mail-flag-data
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/ovi054/mail-flag-data.math-code-qa
Math & Code QA — Instruction Dataset
Worked mathematical solutions and short code answers, built for the
Adaption Labs AutoScientist Challenge (Math & Code category).
Rows
5,200
Math
3,600
Code
1,600
Distinct answers
5,199 (100%)
Duplicate questions
none
Nulls
none
Question length
median 27 words
Answer length
median 58 words (max 89)
License
CC-BY-4.0
What makes the math rows unusual
Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.flan-t5-boosting-bbh_directflask-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Flask
Documentation Data Source Link: https://flask.palletsprojects.com/en/stable/
Data Source License: https://flask.palletsprojects.com/en/stable/license/
Data Source Authors: Pallets Project
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
MNLP_M3_rag_documentsRotBot_Flags
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kellydoesstuff/RotBot_Flags.math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.
