CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roman-bushuiev /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Papers MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.tabularother100K<n<1M23 likes3.5k downloads2mo agoHugging Face02masakhane /afrimgsm Dataset Card for afrimgsm Dataset Summary AFRIMGSM is an evaluation dataset comprising translations of a subset of the GSM8k dataset into 16 African languages. It includes test sets across all 18 languages, maintaining an English and French subsets from the original GSM8k dataset. Languages There are 18 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm.text1K<n<10K10 likes2.3k downloads1y agoHugging Face03masakhane /masakhanews Dataset Card for [Dataset Name] Dataset Summary MasakhaNEWS is the largest publicly available dataset for news topic classification in 16 languages widely spoken in Africa. The train/validation/test sets are available for all the 16 languages. Supported Tasks and Leaderboards [More Information Needed] news topic classification: categorize news articles into new topics e.g business, sport sor politics. Languages There are 16 languages available :… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/masakhanews.texttext-classification10K<n<100K16 likes2k downloads10mo agoHugging Face04masakhane /afrimmlu Dataset Card for afrimmlu Dataset Summary AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages. It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset. Languages There are 17 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.textquestion-answering10K<n<100K12 likes1.7k downloads1y agoHugging Face05masakhane /afrixnli Dataset Card for afrixnli Dataset Summary AFRIXNLI is an evaluation dataset comprising translations of a subset of the XNLI dataset into 16 African languages. It includes both validation and test sets across all 18 languages, maintaining the English and French subsets from the original XNLI dataset. Languages There are 18 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrixnli.texttext-classification10K<n<100K6 likes1.2k downloads2y agoHugging Face06APProjects /us-mass-layoff-and-plant-closing-statistics-monthly US Mass Layoff & Plant Closing Statistics — monthly series, rebuilt every day The federal government stopped publishing this. The Bureau of Labor Statistics ran a Mass Layoff Statistics program until sequestration cut it: the final release was 2013-06-21, covering May 2013 and the program was eliminated on 2013-09-30. Since then there has been no monthly, national, machine-readable count of US mass-layoff and plant-closing events. This series is one answer to that gap, built… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-mass-layoff-and-plant-closing-statistics-monthly.tabulartime-series-forecasting1K<n<10K0 likes686 downloads16h agoHugging Face07masakhane /AfriDocMTdata ├── document │ ├── Health │ │ ├── dev.csv │ │ ├── test.csv │ │ └── train.csv │ └── Tech │ ├── dev.csv │ ├── test.csv │ └── train.csv └── sentence ├── Health │ ├── dev.csv │ ├── test.csv │ └── train.csv └── Tech ├── dev.csv ├── test.csv └── train.csv AFRIDOC-MT is a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/AfriDocMT.texttranslation10K<n<100K9 likes417 downloads1y agoHugging Face08MasahiroKaneko /eagle Eagle 🦅: Ethical Dataset Given from Real Interactions Introduction This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs). If you use the Eagle dataset in your research, please cite the following: @inproceedings{Eagle:arxiv:2024, title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.tabulartext-generation100K<n<1M4 likes390 downloads3y agoHugging Face09APProjects /us-plant-closings-vs-mass-layoffs-warn-act US plant closings vs mass layoffs (WARN Act), normalized The federal WARN Act is a statute about two events: a plant closing and a mass layoff. Every US state publishes its notices with that distinction buried in a free-text column — and across 48 states that column contains 531 distinct raw strings: Closure, Closing *, CL, Plant Closing, facility closure, shutdown operations, LO, WR, Mass Layoff - No Recall, Layoff Permanent, and hundreds more. About a fifth of rows leave it… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-plant-closings-vs-mass-layoffs-warn-act.tabulartabular-classification10K<n<100K0 likes371 downloads16h agoHugging Face10APProjects /massachusetts-layoffs-warn-act-notices-daily Massachusetts WARN Act layoff notices — every filing we hold since 2021, one CSV, rebuilt daily 294 Massachusetts WARN notices — every one this dataset holds, back to 2021 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-31 · state source last checked 2026-09-24T14:04Z · official source: Massachusetts EOLWD (WARN layoff and closure updates) — WARN notices. Massachusetts employers must file a WARN Act notice with the state… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/massachusetts-layoffs-warn-act-notices-daily.tabulartabular-classificationn<1K0 likes354 downloads16h agoHugging Face11tickertruthorg /nse-india-security-master TickerTruth — NSE India Security Master (Explorer) A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline. Why this dataset exists India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruthorg/nse-india-security-master.textother1K<n<10K0 likes290 downloads2mo agoHugging Face12masakhane /AfriADRtexttext-generation10K<n<100K2 likes222 downloads2y agoHugging Face13masakhane /ntrex_african Dataset Summary Multilingual News Test References for MT Evaluation from English into 32 target African languages in the NTREX-128 dataset. Adapted from [NTREX]{https://github.com/MicrosoftTranslator/NTREX/tree/main}, processed into tab-separated value (TSV) files for easy integration into evaluation workflows Sample usage from datasets import load_dataset # Load a specific language configuration dataset = load_dataset("masakhane/ntrex_african", name="afr_Latn"… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/ntrex_african.texttranslation10K<n<100K5 likes216 downloads2y agoHugging Face14tickertruth /nse-india-security-master TickerTruth — NSE India Security Master (Explorer) A clean, normalized reference table of 2,389 NSE-listed equities — ISIN mappings, listing dates, company names, and active/delisted status — built from the TickerTruth India reference-data pipeline. Why this dataset exists India equity data is notoriously messy. NSE symbols change (renames, mergers, delistings), ISINs get reissued, and raw bhavcopy files carry no historical context. TickerTruth's pipeline… See the full description on the dataset page: https://huggingface.co/datasets/tickertruth/nse-india-security-master.textother1K<n<10K0 likes185 downloads4mo agoHugging Face15k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes181 downloads3mo agoHugging Face16ChaseLabs /Harmful-Texts-On-Mastodon 🦣 Mastodon Wild Data for Harmful Content Detection Overview The Harmful Texts on Mastodon dataset is a human-annotated corpus of 3,000 English posts collected from the decentralized social media platform Mastodon between December 2024 and February 2025.It is designed to evaluate the robustness, generalization, and personalization capabilities of large language models (LLMs) and in-context learning (ICL) approaches for harmful content detection in real-world scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/ChaseLabs/Harmful-Texts-On-Mastodon.texttext-classification1K<n<10K2 likes176 downloads11mo agoHugging Face17letrinhan /vn-provinces-admin-land-master Vietnam administrative units and land-use master panel by locality Wide geo×year master joining NSO Đơn vị hành chính and Đất đai locality packs: administrative unit counts, land-use area, land-use structure shares, and natural land-area change indexes. Outer join on geo_code×year. Climate/hydrology station tables (sunshine, rainfall, humidity, temperature, river and sea levels) are out of scope. Province names follow ar_core.vn_geo. Figures Hero Comparison… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-admin-land-master.tabular1K<n<10K0 likes138 downloads5d agoHugging Face18masakhane /afrimgsm-translate-test Dataset Card for afrimgsm-translate-test Dataset Summary AFRIMGSM-TT is an evaluation dataset comprising translations of the GSM8k dataset from 16 African languages and 1 high resource language into English using NLLB. It includes test sets across all 17 languages. Languages There are 17 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm-translate-test.tabulartext-classification1K<n<10K1 likes135 downloads2y agoHugging Face19mastergokul /project-madurai-booksProject Madurai Books Text Dataset This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website. Dataset Details You can see a table above called "Meta Data", which is just an info table. You can't able to preview the "Source Data" table, due to it being about 300MB. [Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.tabulartext-classification1K<n<10K0 likes126 downloads2y agoHugging Face20letrinhan /vn-provinces-society-environment-master Vietnam health, living standards, culture, justice and environment master Wide geo×year master joining 39 NSO locality packs spanning health facilities and workforce, immunization/malnutrition/HIV, culture (libraries, press, heritage), income/HDI/Gini/poverty, water-sanitation-ICT-housing living standards, criminal justice and civil enforcement, traffic/fire incidents, and solid/hazardous waste plus industrial-cluster wastewater treatment. Outer join on geo_code×year.… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-society-environment-master.tabular1K<n<10K0 likes124 downloads5d agoHugging Face21letrinhan /vn-par-master Vietnam PAR Index master panel (province + ministry, 2012-2025) Stacked master of Vietnam's Public Administration Reform Index (PAR INDEX / Chỉ số CCHC) joining the provincial UBND panel and the ministry / ministerial- level agency panel for 2012-2025. Rows are discriminated by unit_type (province | ministry). this is not a geo×agency cross join. Shared score columns: score_moha, score_survey, par_index, par_rank. Leaf packs hold year extracts. this pack ships clean stacked… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-par-master.tabular1K<n<10K1 likes121 downloads5d agoHugging Face22masakhane /afrimmlu-translate-test Dataset Card for afrimmlu-translate-test Dataset Summary AFRIMMLU-TT is an evaluation dataset comprising translations of the AFRIMMLU dataset from 16 African languages and 1 high resource language into English using NLLB. It includes test sets across all 17 languages. Languages There are 17 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu-translate-test.texttext-classification1K<n<10K2 likes116 downloads2y agoHugging Face23rsaunders /tradedatahub-masked-preview-sample TradeDataHub Masked Contractor Preview Sample This is a sample/teaser dataset published by TradeDataHub, a provider of downloadable U.S. contractor datasets organized by state, trade, and city. What this sample contains 12,952 masked preview rows corresponding to TradeDataHub's city/trade products Columns: product_id, business, city, trade, phone_available, website_available, verification_date Business identities are masked ("Masked business") by design: this… See the full description on the dataset page: https://huggingface.co/datasets/rsaunders/tradedatahub-masked-preview-sample.text10K<n<100K0 likes113 downloads18d agoHugging Face24letrinhan /vn-provinces-population-master Vietnam population master panel by locality Wide geo×year master joining all published NSO "Dân số" locality packs (area/density, average population by sex/residence, sex ratio, vital rates, fertility, child mortality, growth, migration, life expectancy, literacy, marriages, age at first marriage, divorces, birth registration, death registrations). Outer join on geo_code×year for 1995-2024. cells are missing where a source pack has no year. Includes historical Ha Tay when… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-population-master.tabular1K<n<10K0 likes110 downloads5d agoHugging Face25letrinhan /vn-provinces-education-master Vietnam education master panel by locality Wide geo×year master joining 16 NSO Giáo dục locality packs: preschool, general schools/classes/teachers/pupils (incl. female and ethnic-minority), pupils per class/teacher, solid classroom rate, upper-secondary graduation, university lecturers/students, and 2023 vocational education. Outer join on geo_code×year (2001-2024). R&D / patents / national ownership tables are out of scope. Province names follow ar_core.vn_geo.… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-education-master.tabular1K<n<10K0 likes109 downloads5d agoHugging Face26letrinhan /vn-provinces-enterprise-master Vietnam enterprise master panel by locality Wide geo×year master joining 20 NSO Doanh nghiệp locality packs: new registrations, operating enterprises (levels and per 1000 population), enterprises with business results, employment (total/female), capital, fixed assets, net revenue, employment- and capital-size bins, labor income, profit, cooperatives, and nonfarm individual establishments. Outer join on geo_code×year (2010-2024). Sector/ownership/technology tables without địa… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-enterprise-master.tabularn<1K0 likes108 downloads5d agoHugging Face27Feng613 /MASS-EX MASS-EX: Expert-Annotated Dataset for Interpretable Sleep Staging 中文版 Associated Paper:Guifeng Deng, Pan Wang, Jiquan Wang, Shuying Rao, Junyi Xie, Wanjun Guo, Tao Li, Haiteng Jiang. "SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model." arXiv preprint, 2026. arXiv:2603.26738 Authors Name Affiliation ORCID Guifeng Deng Zhejiang University 0009-0001-1940-7797 Pan Wang Wenzhou Medical University 0009-0001-6664-6934 Wanjun… See the full description on the dataset page: https://huggingface.co/datasets/Feng613/MASS-EX.texttext-classification10K<n<100K0 likes106 downloads6mo agoHugging Face28letrinhan /vn-provinces-trade-tourism-transport-master Vietnam trade, tourism and transport master panel by locality Wide geo×year master joining 15 NSO Thương mại / Du lịch / Vận tải locality packs: retail trade and service revenue, markets, supermarkets, shopping centers, travel-agency revenue, and passenger/freight transport volume and turnover (total, road, inland waterway). Outer join on geo_code×year (1995-2024). Merchandise trade, visitor expenditure, ports/air, and post/telecom national series are out of scope. Province… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-trade-tourism-transport-master.tabular1K<n<10K0 likes106 downloads5d agoHugging Face29mashnoor /hls-ast-sagehls SAGE-HLS Dataset: AST-Guided HLS-C Code Generation The SAGE-HLS dataset is a large-scale, synthesis-friendly dataset for natural language to HLS-C code generation, enhanced by AST (Abstract Syntax Tree) representations. It supports training LLMs to generate high-quality, synthesis-ready high-level synthesis (HLS) code from functional descriptions, with structure-aware guidance. 📦 Dataset Structure Each sample in the dataset contains the following fields: Field… See the full description on the dataset page: https://huggingface.co/datasets/mashnoor/hls-ast-sagehls.texttext-generation10K<n<100K0 likes102 downloads1y agoHugging Face30letrinhan /vn-provinces-investment-construction-master Vietnam investment and construction master panel by locality Wide geo×year master joining NSO Đầu tư và Xây dựng locality packs: FDI licensed stock as of end-2024, FDI licensed flow in 2024, completed housing floor area, and self-built housing floor area. Outer join on geo_code×year (2010-2024). FDI columns are populated for 2024 only. Society-wide investment by ownership/sector, outbound FDI, and 2019 social-housing-by-region tables are out of scope. Province names follow… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-investment-construction-master.tabular1K<n<10K0 likes101 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.