datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.usgs-global-earthquake-catalog
USGS Global Earthquake Catalog
Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service.
Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including:
Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event.
Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.ecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.sms_spam_categorymorphscore
MorphScore
MorphScore is a tokenizer evaluation framework, which evaluates the extent to which a tokenizer segments words along morpheme boundaries. This repository contains the datasets used to calculate MorphScore.
In total, we have datasetes for 86 languages, but only 70 languages have at least 100 items after filtering.
All datasets are derived from existing Universal Dependencies treebanks. In the table below, we link the source dataset for each language.
See the new preprint… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/morphscore.BATIS
license: cc-by-nc-4.0
BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models
This repository contains the dataset used in experiments shown in BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models. To download the dataset, you can use the load_dataset function from HuggingFace. For example :
from datasets import load_dataset
# Training Split for Kenya
training_kenya = load_dataset("cathv/BATIS", name="Kenya"… See the full description on the dataset page: https://huggingface.co/datasets/cathv/BATIS.shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.us-bank-transaction-categories-v2
US Bank Transaction Categories v2 — Synthetic Dataset
68,000 sign-prefixed transaction descriptions across 17 spending categories, modeled after real US bank statement formats. Designed for training classifiers that work on actual bank data — not the clean "Starbucks coffee" descriptions that most datasets use.
Successor to v1.
Why This Dataset
Real bank transaction data is private. But the formats are universal — Chase, Apple Card, PayPal, Capital One, Mercury all… See the full description on the dataset page: https://huggingface.co/datasets/DoDataThings/us-bank-transaction-categories-v2.furniture-catalog
Furniture Catalog
Image catalog and training splits from the bachelor's thesis "Visual Furnishings Compatibility Learning and Retrieval Using Machine Learning" (Ukrainian Catholic University, 2026).
5 171 individual furniture images across two room types, 1 781 room scene images, plus triplet training data (golden / train / val splits) used to train the compatibility model.
Repo layout
{room}/
{category}/ *.jpg — individual furniture images… See the full description on the dataset page: https://huggingface.co/datasets/Darebal/furniture-catalog.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.PoseX
PoseX: AI Defeats Physics Approaches on Protein-Ligand Cross Docking
Paper | GitHub | Leaderboard
PoseX is a comprehensive benchmark dataset designed to evaluate molecular docking algorithms for predicting protein-ligand binding poses. It includes curated datasets for both self-docking and cross-docking scenarios.
The file posex_set.zip contains the processed dataset for docking evaluation, while posex_cif.zip contains the raw CIF files from RCSB PDB. For information about creating… See the full description on the dataset page: https://huggingface.co/datasets/CataAI/PoseX.narit-ghosts-halo-catalogs
NARIT GHOSTS Halo Catalogs
Dataset Description
This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST).
It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.CatCoLA
GENERAL INFORMATION
1. Dataset title:
CatCoLA - Catalan Corpus of Linguistic Acceptability
2. Authorship:
Name: Núria Bel
Institution: Universitat Pompeu Fabra, UPF
Email: nuria.bel@upf.edu
ORCID: 0000-0001-9346-7803
Name: Marta Punsola
Institution: Universitat Pompeu Fabra, UPF
Email: marta.punsola@gmail.com
ORCID:
Name: Valle Ruiz-Fernández
Institution: Barcelona Supercomputing Center
Email: valle.ruizfernández@bsc.es
ORCID:
DESCRIPTION… See the full description on the dataset page: https://huggingface.co/datasets/nbel/CatCoLA.USC-Course-Catalog
USC Course Catalog - 2024, Spring Term
This dataset consists of all classes provided by USC (as of December 1, 2023) that USC is providing in 2024 Spring.
While it is a small dataset, this could be used in some finetuning, generation, or RAG application tasks. One example would be this -> https://huggingface.co/spaces/USC/USC-GPT
I will also be scraping the 2024 fall term classes when they are released by USC!
If you want the web scraping script I used for this, feel free to send me… See the full description on the dataset page: https://huggingface.co/datasets/USC/USC-Course-Catalog.Yahoo_Answers_10_categories_for_NLP
Dataset Card for Dataset Name
The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information.
Dataset Description
The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.online_shopping_10_catsbiomedical-topic-categorizationvn-provinces-cattle-count
Vietnam provinces cattle count
Number of cattle (bo) (thousand heads). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1866 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-cattle-count.shopify-products-catalog-sample
Shopify Products Catalog with Price History (DTC Stores) — Free Sample
A free sample of a Shopify e-commerce catalog dataset built from public, login-free products.json endpoints of independent DTC (direct-to-consumer) Shopify stores.
This sample build covers 386 stores / 533,479 SKU rows (2026-08-01 snapshot). Full-run scale on record: 4,100+ stores / 6.18M SKUs, refreshed with weekly snapshots that also yield per-variant price-change and availability-change events.
➡️ Full… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/shopify-products-catalog-sample.News_Articles_Categorization
Dataset Card for News_Articles_Categorization
Dataset Description
3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Text and Category.
The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.Catacomp-104
CataCompDetect
Video dataset of cataract surgeries labeled for surgical complications
(Posterior Capsule Rupture, Vitreous Loss, Iris Prolapse), used to train and
evaluate the CataCompDetect
complication-detection pipeline.
Dataset structure
catacomp-hf/
├── train/
│ ├── metadata.csv
│ ├── catacomp-train-001.mp4
│ ├── catacomp-train-002.mp4
│ └── ...
└── val/
├── metadata.csv
├── catacomp-val-001.mp4
├── catacomp-val-002.mp4
└── ...
Videos… See the full description on the dataset page: https://huggingface.co/datasets/SankaraEyeHospital/Catacomp-104.CatPred-DB
CatPred-DB: Enzyme Kinetic Parameters Database
Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters
GitHub: https://github.com/maranasgroup/CatPred-DB
Dataset Description
CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters:
Parameter
Description
Datapoints
kcat
Turnover… See the full description on the dataset page: https://huggingface.co/datasets/kunikohunter/CatPred-DB.Politifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
vn-provinces-cattle-meat-production
Vietnam provinces cattle meat production
Cattle meat production (liveweight) (thousand tons). Coverage 2018-2024. Year 2024 is preliminary. Source values in tons are converted to thousand tons. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-cattle-meat-production.french_narrativeqa
Description
Dataframe containing 143 French books in txt format.More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day)
the question column contains a single question asked about the associated text
the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.indian-transaction-categorization-synthetic
Synthetic Indian Bank Transaction Narrations
810 synthetic (text, category) pairs mimicking Indian bank/credit-card statement narrations —
built to train the Sumeetgpt/indian-transaction-categorizer
SetFit model.
Why this exists
While building a personal finance app, we searched for a public dataset pairing real Indian
transaction narration formats (UPI, NEFT, IMPS, ACH) with spending-category labels, and found
none: datasets with real-looking Indian narration… See the full description on the dataset page: https://huggingface.co/datasets/Sumeetgpt/indian-transaction-categorization-synthetic.Sinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc.
If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.Aspect-Based_Sentiment_Analysis_for_Catering
说明
数据集来源于AI Challenger 2018
sentiment_analysis_trainingset.csv 为训练集数据文件,共105000条评论数据
sentiment_analysis_validationset.csv 为验证集数据文件,共15000条评论数据
sentiment_analysis_testa.csv 为测试集A数据文件,共15000条评论数据
数据集分为训练、验证、测试A与测试B四部分。数据集中的评价对象按照粒度不同划分为两个层次,层次一为粗粒度的评价对象,例如评论文本中涉及的服务、位置等要素;层次二为细粒度的情感对象,例如“服务”属性中的“服务人员态度”、“排队等候时间”等细粒度要素。评价对象的具体划分如下表所示。
The dataset is divided into four parts: training, validation, test A and test B. This dataset builds a two-layer labeling system according to the… See the full description on the dataset page: https://huggingface.co/datasets/xcz0/Aspect-Based_Sentiment_Analysis_for_Catering.libero-plus-episode-categories
LIBERO-Plus episode → perturbation-category labels
Which perturbation category each of the 14,347 LIBERO-Plus episodes belongs to.
LIBERO-Plus perturbs a base LIBERO task along several axes. The release these
labels came from contains five of them --- the project describes more, so treat
this as the taxonomy of this 14,347-episode release, not of LIBERO-Plus as a
whole. The
lerobot/libero_plus conversion does not carry that label, so you cannot ask
"how does my policy do under… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/libero-plus-episode-categories.
