datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huggingface-spaces-codes
📊 Dataset Description
This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data.
📝 Data Fields
Field
Type
Description
repository
string
Huggingface Spaces repository names.
sdk
string
Software Development Kit of the space.
license
string
License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.SpatialVIDSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/SpatialVID/SpatialVID.SpatialVID-HQSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/FelixYuan/SpatialVID-HQ.Sparkle
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.sms-spam-collection
SMS Spam Collection v.1
DESCRIPTION
The SMS Spam Collection v.1 (hereafter the corpus) is a set of SMS tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged acording being ham (legitimate) or spam.
1.1. Compilation
This corpus has been collected from free or free for research sources at the Web:
A collection of between 425 SMS spam messages extracted manually from the Grumbletext Web… See the full description on the dataset page: https://huggingface.co/datasets/codesignal/sms-spam-collection.spacetravlr
SpaceTravLR dataset hub
Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports.
Layout
spacetravlr/
├── tonsil/ # placeholder / demo gene outputs
└── xenium_skin_mixed/
├── run.toml # shared training config for this cohort
├── manifest.json # sample index and upload metadata
├── sample12/
├── sample13/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.us-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.fsrs-datasetspambase
Spambase
The Spambase dataset from the UCI ML repository.
Is the given mail spam?
Configurations and tasks
Configuration
Task
Description
spambase
Binary classification
Is the mail spam?
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/spambase")["train"]
aopsSpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.SpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.spacr_settingsthomas-2018-spark-wt
SPARK (wild-type accumulator phenotype): Human-curated and standardized MICs
These data were collated by the authors of:
Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell
Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery
ACS Infectious Diseases 2018 4 (11), 1536-1539
DOI: 10.1021/acsinfecdis.8b00193
We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values,
give succint column… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-wt.Spatial-DISE
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
📋 Overview
Spatial-DISE is a comprehensive benchmark dataset designed to evaluate spatial reasoning capabilities in vision-language models. The dataset focuses on various aspects of spatial intelligence including 3D perception, spatial transformation, and geometric reasoning across multiple difficulty levels.
🧪 Evaluation Support
Supported:… See the full description on the dataset page: https://huggingface.co/datasets/TACPS-liv/Spatial-DISE.spatial457_mcqSMS-spamThis dataset can be found on Kaggle, Huggingface and many other websites, but the source is from the research paper [Tiago], whose authors contributed it to
the Machine Learning repository at [UCI]. It contains 5,574 SMS messages, of which 747 massages are labeled as spam.
By nature, the messages are short and, in some cases, quite cryptic and personal.
The CSV file is a straightforward representation of the data.
References
[UCI] https://archive.ics.uci.edu/dataset/228/sms+spam+collection… See the full description on the dataset page: https://huggingface.co/datasets/bvk/SMS-spam.SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.sms-spam-classificationSpam_SMS
Description
The Spam SMS is a set of SMS-tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam.
Source: uciml/sms-spam-collection-dataset
energy-consumption-hourly-spainsms_spam_collectionsms-otp-spam-dataset
📲 SMS OTP Spam Dataset
A synthetic dataset of 10,000 OTP-style SMS messages for spam classification tasks. The dataset includes both valid and spam-like messages, with labels for message validity and delivery status.
📊 Dataset Summary
Total samples: 10,000
Valid messages: 90%
Not valid (spam-like): 10%
Status types: delivered, failed, spam, bounced, expired
Each entry includes:
phone_id: Synthetic phone number
sms_text: Message content
label: valid or not valid… See the full description on the dataset page: https://huggingface.co/datasets/alusci/sms-otp-spam-dataset.spaces-of-the-week-legacysms_spam_categorySMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.punta-cana-spanish-reviewsThis data set was collected for academic purposes, suitable for some NLP tasks including sentiment analysis.
FSRS-Anki-20k
Update
We have released a new dataset: anki-revlogs-10k.
Introduction
FSRS-Anki-20k is a dataset of 20k collections from Anki for FSRS project. It is a random sample of collections with 5000+ revlog entries, so it should contain a mix of older (still active) users, and newer users. Entries are pre-sorted in (cid, id) order.
There are two versions of the dataset: ./revlogs and ./dataset. The ./revlogs version contains the raw revlog entries, while the ./dataset version… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/FSRS-Anki-20k.spam-messages
Dataset
The dataset is composed of messages labeled by ham or spam, merged from three data sources:
SMS Spam Collection https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset
Telegram Spam Ham https://huggingface.co/datasets/thehamkercat/telegram-spam-ham/tree/main
Enron Spam: https://huggingface.co/datasets/SetFit/enron_spam/tree/main (only used message column and labels)
The prepare script for enron is available at… See the full description on the dataset page: https://huggingface.co/datasets/mshenoda/spam-messages.
