datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.nba-games
NBA Games Data
This data is an updated version of the original NBA
Games by Nathan Lauga.
Data source
Code
Updated to: 2025-02-13
The dataset retains the original format and includes the following files:
games.csv – Summary of NBA games, including scores and team details.
games_details.csv – Detailed player statistics for each game.
players.csv – Player information.
ranking.csv – Daily NBA team rankings.
teams.csv – List of all NBA teams.
skin-cancer-ham10000-datasetIntentQAham10ktelegram-spam-hamnlp_twitter_analysisthe-biggest-spam-ham-phish-email-dataset-300000
The Biggest Spam Ham Phish Email Dataset (250000+)
This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects.
The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.ham10000-skin-lesion-classifier-datalegal_summarycode-switching-codesaviours-si26-hamzaBBC-SinhalaCNN-Daily-Mail-Sinhala
Dataset Summary
This dataset card aims to be creating a new dataset or Sinhala news summarization tasks. It has been generated using [https://huggingface.co/datasets/cnn_dailymail] and google translate.
Data Instances
For each instance, there is a string for the article, a string for the highlights, and a string for the id. See the CNN / Daily Mail dataset viewer to explore more examples.
{'id': '0054d6d30dbcad772e20b22771153a2a9cbeaf62',
'article': '(CNN) -- An American… See the full description on the dataset page: https://huggingface.co/datasets/Hamza-Ziyard/CNN-Daily-Mail-Sinhala.nlp-assignment-news-data
Dataset Structure
Data Instances
{
"headline": "string",
"label": "string"
}
Data Fields
The data fields are:
headline: a string feature.
label: a classification label, with possible values including positive, negative and neutral.
Data Splits
The nlp-assignment-news-data dataset has 3 splits: train, validation, and test.
How to use it
from datasets import load_dataset
# This download train, validation and test sets.
ds =… See the full description on the dataset page: https://huggingface.co/datasets/hamza-student-123/nlp-assignment-news-data.spam_ham_commentshamburger-bun-prices-raw-dataset-2026
24,551 raw U.S. hamburger bun price observations across 12 ZIP markets and 29 days.
Hamburger Bun Prices Raw Dataset (2026)
Analyze 24,551 unaggregated product-level listed retail prices for hamburger buns across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated product-level observations… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/hamburger-bun-prices-raw-dataset-2026.mail_spam_ham_dataset
Mail dataset(spam and ham)
5616 rows
sample-OpenOrcaAshrafur_bangla_mathSPAM_HAM_DATASETmyanmar-township-boundaries-mimu
Myanmar Township Boundaries (MIMU v9.4)
This repository provides documentation, metadata, and reference structures for the Myanmar Township Boundaries (Administrative Level 3) dataset, based on the Myanmar Information Management Unit (MIMU) Pcode version 9.4.
Dataset Overview
Source Organization: Myanmar Information Management Unit (MIMU-GIS)
Original Publication Date: June 18, 2023
Edition/Version: Pcode v9.4
Administrative Level: Admin3 (Township)
Scale: Digitized at… See the full description on the dataset page: https://huggingface.co/datasets/hamiwirrr/myanmar-township-boundaries-mimu.iranian-social-norms-dataset
Iranian Social Norms (ISN) Dataset
Overview
The Iranian Social Norms (ISN) dataset is a curated collection of social norms specific to Iranian culture. It provides a rich resource for studying cultural norms and their variations across different demographic groups. By incorporating demographic context, such as age, gender, ethnicity, and religion, this dataset enables a nuanced exploration of social expectations within Iranian society.
Dataset Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/hamiids/iranian-social-norms-dataset.arabic_dialectsall-restaurants-in-manchester-new-hampshire-us-135738
All Restaurants in Manchester, New Hampshire, US
Free sample dataset from BeamStation
This dataset contains a complete export of every restaurant operating in Manchester, New Hampshire, US, with 537 records refreshed on a weekly basis. Each record includes the full profile of an establishment—name, address, contact information, cuisine type, hours of operation, and any other columns made available by the source. It is ideal for developers building local‑search applications, market… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-manchester-new-hampshire-us-135738.pybamm-docsurdu-spam-dataset
Urdu Spam Detection Dataset
Description
This dataset is designed for classifying Urdu text into:
0 → Not Spam
1 → Spam
It is intended for AI-powered emergency helpline systems (e.g., 1122/911) to filter prank or irrelevant calls.
Dataset Structure
Format: CSV
Column
Type
Description
text
string
Urdu sentence
label
int (0/1)
Spam classification
Example
text,label
آپ کو میں نے پہلے بھی کال کیا تھا کیا یاد ہے,1
یہ ایک ایمرجنسی ہے… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-spam-dataset.CosFunc-HO
CosFunc-HO: A Hierarchical Cosmetic Ingredient Function Dataset
Overview
CosFunc-HO is a cleaned and structured cosmetic ingredient function dataset developed from COSDNA data for research in Large Language Models (LLMs), ontology learning, multi-label classification, and cosmetic chemistry analysis.
The dataset contains more than 9,000 labeled cosmetic ingredient records and includes hierarchical functional annotations organized into multiple levels of abstraction.
This… See the full description on the dataset page: https://huggingface.co/datasets/hamdanzameer/CosFunc-HO.GalaxyMergerTabularTensorSamplemerged_characters_tinyllamaTCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/hammad655/TCGA-Cancer-Variant-and-Clinical-Data.
