datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Papers
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link
MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.afrimgsm
Dataset Card for afrimgsm
Dataset Summary
AFRIMGSM is an evaluation dataset comprising translations of a subset of the GSM8k dataset into 16 African languages.
It includes test sets across all 18 languages, maintaining an English and French subsets from the original GSM8k dataset.
Languages
There are 18 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm.masakhanews
Dataset Card for [Dataset Name]
Dataset Summary
MasakhaNEWS is the largest publicly available dataset for news topic classification in 16 languages widely spoken in Africa.
The train/validation/test sets are available for all the 16 languages.
Supported Tasks and Leaderboards
[More Information Needed]
news topic classification: categorize news articles into new topics e.g business, sport sor politics.
Languages
There are 16 languages available :… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/masakhanews.afrimmlu
Dataset Card for afrimmlu
Dataset Summary
AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages.
It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.afrixnli
Dataset Card for afrixnli
Dataset Summary
AFRIXNLI is an evaluation dataset comprising translations of a subset of the XNLI dataset into 16 African languages.
It includes both validation and test sets across all 18 languages, maintaining the English and French subsets from the original XNLI dataset.
Languages
There are 18 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrixnli.us-mass-layoff-and-plant-closing-statistics-monthly
US Mass Layoff & Plant Closing Statistics — monthly series, rebuilt every day
The federal government stopped publishing this. The Bureau of Labor
Statistics ran a Mass Layoff Statistics program until sequestration cut it:
the final release was 2013-06-21, covering May 2013
and the program was eliminated on 2013-09-30.
Since then there has been no monthly, national, machine-readable count of US
mass-layoff and plant-closing events.
This series is one answer to that gap, built… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-mass-layoff-and-plant-closing-statistics-monthly.AfriDocMTdata
├── document
│ ├── Health
│ │ ├── dev.csv
│ │ ├── test.csv
│ │ └── train.csv
│ └── Tech
│ ├── dev.csv
│ ├── test.csv
│ └── train.csv
└── sentence
├── Health
│ ├── dev.csv
│ ├── test.csv
│ └── train.csv
└── Tech
├── dev.csv
├── test.csv
└── train.csv
AFRIDOC-MT is a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/AfriDocMT.eagle
Eagle 🦅: Ethical Dataset Given from Real Interactions
Introduction
This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs).
If you use the Eagle dataset in your research, please cite the following:
@inproceedings{Eagle:arxiv:2024,
title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.us-plant-closings-vs-mass-layoffs-warn-act
US plant closings vs mass layoffs (WARN Act), normalized
The federal WARN Act is a statute about two events: a plant closing and a
mass layoff. Every US state publishes its notices with that distinction buried in
a free-text column — and across 48 states that column contains
531 distinct raw strings: Closure, Closing *, CL,
Plant Closing, facility closure, shutdown operations, LO, WR,
Mass Layoff - No Recall, Layoff Permanent, and hundreds more. About a fifth of
rows leave it… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-plant-closings-vs-mass-layoffs-warn-act.massachusetts-layoffs-warn-act-notices-daily
Massachusetts WARN Act layoff notices — every filing we hold since 2021, one CSV, rebuilt daily
294 Massachusetts WARN notices — every one this dataset holds, back to 2021 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-31
· state source last checked 2026-09-24T14:04Z · official source: Massachusetts EOLWD (WARN layoff and closure updates) — WARN notices.
Massachusetts employers must file a WARN Act notice with the state… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/massachusetts-layoffs-warn-act-notices-daily.nse-india-security-master
TickerTruth — NSE India Security Master (Explorer)
A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline.
Why this dataset exists
India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruthorg/nse-india-security-master.AfriADRntrex_african
Dataset Summary
Multilingual News Test References for MT Evaluation from English into 32 target African languages in the NTREX-128 dataset. Adapted from [NTREX]{https://github.com/MicrosoftTranslator/NTREX/tree/main}, processed into tab-separated value (TSV) files for easy integration into evaluation workflows
Sample usage
from datasets import load_dataset
# Load a specific language configuration
dataset = load_dataset("masakhane/ntrex_african", name="afr_Latn"… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/ntrex_african.nse-india-security-master
TickerTruth — NSE India Security Master (Explorer)
A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline.
Why this dataset exists
India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruth/nse-india-security-master.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.Harmful-Texts-On-Mastodon
🦣 Mastodon Wild Data for Harmful Content Detection
Overview
The Harmful Texts on Mastodon dataset is a human-annotated corpus of 3,000 English posts collected from the decentralized social media platform Mastodon between December 2024 and February 2025.It is designed to evaluate the robustness, generalization, and personalization capabilities of large language models (LLMs) and in-context learning (ICL) approaches for harmful content detection in real-world scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/ChaseLabs/Harmful-Texts-On-Mastodon.vn-provinces-admin-land-master
Vietnam administrative units and land-use master panel by locality
Wide geo×year master joining NSO Đơn vị hành chính and Đất đai locality packs: administrative unit counts, land-use area, land-use structure shares, and natural land-area change indexes. Outer join on geo_code×year. Climate/hydrology station tables (sunshine, rainfall, humidity, temperature, river and sea levels) are out of scope. Province names follow ar_core.vn_geo.
Figures
Hero
Comparison… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-admin-land-master.afrimgsm-translate-test
Dataset Card for afrimgsm-translate-test
Dataset Summary
AFRIMGSM-TT is an evaluation dataset comprising translations of the GSM8k dataset from 16 African languages and 1 high resource language into English using NLLB.
It includes test sets across all 17 languages.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm-translate-test.project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.vn-provinces-society-environment-master
Vietnam health, living standards, culture, justice and environment master
Wide geo×year master joining 39 NSO locality packs spanning health facilities and workforce, immunization/malnutrition/HIV, culture (libraries, press, heritage), income/HDI/Gini/poverty, water-sanitation-ICT-housing living standards, criminal justice and civil enforcement, traffic/fire incidents, and solid/hazardous waste plus industrial-cluster wastewater treatment. Outer join on geo_code×year.… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-society-environment-master.vn-par-master
Vietnam PAR Index master panel (province + ministry, 2012-2025)
Stacked master of Vietnam's Public Administration Reform Index (PAR INDEX / Chỉ số CCHC) joining the provincial UBND panel and the ministry / ministerial- level agency panel for 2012-2025. Rows are discriminated by unit_type (province | ministry). this is not a geo×agency cross join. Shared score columns: score_moha, score_survey, par_index, par_rank. Leaf packs hold year extracts. this pack ships clean stacked… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-par-master.afrimmlu-translate-test
Dataset Card for afrimmlu-translate-test
Dataset Summary
AFRIMMLU-TT is an evaluation dataset comprising translations of the AFRIMMLU dataset from 16 African languages and 1 high resource language into English using NLLB.
It includes test sets across all 17 languages.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu-translate-test.tradedatahub-masked-preview-sample
TradeDataHub Masked Contractor Preview Sample
This is a sample/teaser dataset published by TradeDataHub, a provider of downloadable U.S. contractor datasets organized by state, trade, and city.
What this sample contains
12,952 masked preview rows corresponding to TradeDataHub's city/trade products
Columns: product_id, business, city, trade, phone_available, website_available, verification_date
Business identities are masked ("Masked business") by design: this… See the full description on the dataset page: https://huggingface.co/datasets/rsaunders/tradedatahub-masked-preview-sample.vn-provinces-population-master
Vietnam population master panel by locality
Wide geo×year master joining all published NSO "Dân số" locality packs (area/density, average population by sex/residence, sex ratio, vital rates, fertility, child mortality, growth, migration, life expectancy, literacy, marriages, age at first marriage, divorces, birth registration, death registrations). Outer join on geo_code×year for 1995-2024. cells are missing where a source pack has no year. Includes historical Ha Tay when… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-population-master.vn-provinces-education-master
Vietnam education master panel by locality
Wide geo×year master joining 16 NSO Giáo dục locality packs: preschool, general schools/classes/teachers/pupils (incl. female and ethnic-minority), pupils per class/teacher, solid classroom rate, upper-secondary graduation, university lecturers/students, and 2023 vocational education. Outer join on geo_code×year (2001-2024). R&D / patents / national ownership tables are out of scope. Province names follow ar_core.vn_geo.… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-education-master.vn-provinces-enterprise-master
Vietnam enterprise master panel by locality
Wide geo×year master joining 20 NSO Doanh nghiệp locality packs: new registrations, operating enterprises (levels and per 1000 population), enterprises with business results, employment (total/female), capital, fixed assets, net revenue, employment- and capital-size bins, labor income, profit, cooperatives, and nonfarm individual establishments. Outer join on geo_code×year (2010-2024). Sector/ownership/technology tables without địa… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-enterprise-master.MASS-EX
MASS-EX: Expert-Annotated Dataset for Interpretable Sleep Staging
中文版
Associated Paper:Guifeng Deng, Pan Wang, Jiquan Wang, Shuying Rao, Junyi Xie, Wanjun Guo, Tao Li, Haiteng Jiang. "SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model." arXiv preprint, 2026. arXiv:2603.26738
Authors
Name
Affiliation
ORCID
Guifeng Deng
Zhejiang University
0009-0001-1940-7797
Pan Wang
Wenzhou Medical University
0009-0001-6664-6934
Wanjun… See the full description on the dataset page: https://huggingface.co/datasets/Feng613/MASS-EX.vn-provinces-trade-tourism-transport-master
Vietnam trade, tourism and transport master panel by locality
Wide geo×year master joining 15 NSO Thương mại / Du lịch / Vận tải locality packs: retail trade and service revenue, markets, supermarkets, shopping centers, travel-agency revenue, and passenger/freight transport volume and turnover (total, road, inland waterway). Outer join on geo_code×year (1995-2024). Merchandise trade, visitor expenditure, ports/air, and post/telecom national series are out of scope. Province… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-trade-tourism-transport-master.hls-ast-sagehls
SAGE-HLS Dataset: AST-Guided HLS-C Code Generation
The SAGE-HLS dataset is a large-scale, synthesis-friendly dataset for natural language to HLS-C code generation, enhanced by AST (Abstract Syntax Tree) representations. It supports training LLMs to generate high-quality, synthesis-ready high-level synthesis (HLS) code from functional descriptions, with structure-aware guidance.
📦 Dataset Structure
Each sample in the dataset contains the following fields:
Field… See the full description on the dataset page: https://huggingface.co/datasets/mashnoor/hls-ast-sagehls.vn-provinces-investment-construction-master
Vietnam investment and construction master panel by locality
Wide geo×year master joining NSO Đầu tư và Xây dựng locality packs: FDI licensed stock as of end-2024, FDI licensed flow in 2024, completed housing floor area, and self-built housing floor area. Outer join on geo_code×year (2010-2024). FDI columns are populated for 2024 only. Society-wide investment by ownership/sector, outbound FDI, and 2019 social-housing-by-region tables are out of scope. Province names follow… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-investment-construction-master.
