CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes45k downloads1y agoHugging Face02Weyaxi /huggingface-spaces-codes 📊 Dataset Description This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data. 📝 Data Fields Field Type Description repository string Huggingface Spaces repository names. sdk string Software Development Kit of the space. license string License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.text10K<n<100K12 likes23k downloads3y agoHugging Face03albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads3d agoHugging Face04ZeroOneCreative /amara-spatial-10k AmaraSpatial-10K A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing 10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines. Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.imagetext-to-3d10K<n<100K11 likes11k downloads5mo agoHugging Face05EasonXiao-888 /SpatialEdit-500K SpatialEdit-500K SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware edits such as object moving, object rotation, and camera viewpoint change. The dataset was introduced in the paper SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing. It is generated with a controllable rendering pipeline to provide structured spatial transformations at scale. Project Resources GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/EasonXiao-888/SpatialEdit-500K.imageimage-to-image100K<n<1M14 likes11k downloads6mo agoHugging Face06Spawning /pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset is compatible with webdataset. It was made public after obtaining permission from the original authors of the dataset. You can use the following to explore the dataset with webdataset: import webdataset as wds dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar" dataset = ( wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.image10M<n<100M21 likes10k downloads2y agoHugging Face07spatial-reason /qwen_trajectories_finalimage1K<n<10K0 likes7.8k downloads6mo agoHugging Face08DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes7.1k downloads16d agoHugging Face09ucirvine /sms_spam Dataset Card for [Dataset Name] Dataset Summary The SMS Spam Collection v.1 is a public set of SMS labeled messages that have been collected for mobile phone spam research. It has one collection composed by 5,574 English, real and non-enconded messages, tagged according being legitimate (ham) or spam. Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure Data Instances [More Information… See the full description on the dataset page: https://huggingface.co/datasets/ucirvine/sms_spam.texttext-classification1K<n<10K58 likes6.4k downloads2y agoHugging Face10theislab /SpatialCorpus-110Mtextn<1K17 likes6.2k downloads1y agoHugging Face11OpenDriveLab /SparseVideoNav SparseVideoNav Datasets This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav: BVN: Beyond-the-View Navigation. IFN: Instruction-Following Navigation. Project links: Project page: https://opendrivelab.com/SparseVideoNav GitHub: https://github.com/OpenDriveLab/SparseVideoNav Paper: https://arxiv.org/abs/2602.05827 Dataset Summary SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.tabularrobotics10K<n<100K4 likes6.1k downloads27d agoHugging Face12msr-spare-1 /nemotron-3-nano-30b-20260719-spare-games-envs Nemotron-3-Nano-30B SPARE Self-Play Environments (run_20260719_final) This dataset packages the self-play generated game environments produced by a live SPARE (Self-Play with Adaptive cuRriculum Extension) training run of NVIDIA-Nemotron-3-Nano-30B-A3B. It is a raw-data export for another agent to pick up, replay, and build its own visualization / weave log from. Provenance Run: run_20260719_final Source Ray job: spare_nemotron_games_mtpg768_1784556397 (the live… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/nemotron-3-nano-30b-20260719-spare-games-envs.textn<1K0 likes6k downloads2mo agoHugging Face13SpatialVID /SpatialVIDgatedSpatialVID: A Large-Scale Video Dataset with Spatial Annotations Jiahao Wang1*  Yufeng Yuan1*  Rujie Zheng1*  Youtian Lin1  Jian Gao1  Lin-Zhuo Chen1  Yajie Bao1  Yi Zhang1  Chang Zeng1  Yanxi Zhou1  Xiaoxiao Long1  Hao Zhu1  Zhaoxiang Zhang2  Xun Cao1  Yao Yao1† 1Nanjing University  2Institute of Automation, Chinese Academy of Science  *Equal Contribution  †Corresponding Author CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/SpatialVID/SpatialVID.tabulartext-to-video1M<n<10M47 likes6k downloads6mo agoHugging Face14SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes5.5k downloads5y agoHugging Face15FelixYuan /SpatialVID-HQgatedSpatialVID: A Large-Scale Video Dataset with Spatial Annotations Jiahao Wang1*  Yufeng Yuan1*  Rujie Zheng1*  Youtian Lin1  Jian Gao1  Lin-Zhuo Chen1  Yajie Bao1  Yi Zhang1  Chang Zeng1  Yanxi Zhou1  Xiaoxiao Long1  Hao Zhu1  Zhaoxiang Zhang2  Xun Cao1  Yao Yao1† 1Nanjing University  2Institute of Automation, Chinese Academy of Science  *Equal Contribution  †Corresponding Author CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/FelixYuan/SpatialVID-HQ.tabulartext-to-video100K<n<1M33 likes5.4k downloads6mo agoHugging Face16Biomedical-TeMU /SPACCC_Tokenizer The Tokenizer for Clinical Cases Written in Spanish Introduction This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish. This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.text10K<n<100K0 likes4.5k downloads5y agoHugging Face17stdKonjac /Sparkle Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.imagetext-to-video100K<n<1M1 likes4.4k downloads5mo agoHugging Face18LLDDSS /Awesome_Spatial_VQA_Benchmarksimage10K<n<100K1 likes4.1k downloads1y agoHugging Face19Spawning /PD12M PD12M Summary At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time. Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline Paper Datasheet Project About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.image10M<n<100M186 likes3.4k downloads2y agoHugging Face20spaicom-lab /semasia-mnist Latents for mnist (timm) &nbsp;&nbsp;&nbsp; This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on mnist, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability. Each config corresponds to a single model; only that model's Parquet files are read on load_dataset. Usage Load with datasets and convert to… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-mnist.tabularfeature-extraction100M<n<1B0 likes3.1k downloads3mo agoHugging Face21juliensimon /space-track-tle-history Space-Track TLE History Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports. Quick Start from datasets import load_dataset # Load a specific year ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet") # Load everything (238M rows — use streaming for large-scale analysis) ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.tabulartime-series-forecasting100M<n<1B2 likes3.1k downloads10h agoHugging Face22spanish-ir /messirveJuly 2025 UPDATE: We released version 1.1, adding almost 200k new queries 🎉🎉🎉. v1.2 further adds the article titles as columns for convenience. Use with: country = "full" # "ar", "bo", ... version = "1.2" dataset = datasets.load_dataset("spanish-ir/messirve", country, revision=version) print(dataset) Dataset Card for MessIRve MessIRve is a large-scale dataset for Spanish IR, designed to better capture the information needs of Spanish speakers across different countries.… See the full description on the dataset page: https://huggingface.co/datasets/spanish-ir/messirve.tabulartext-retrieval1M<n<10M21 likes3.1k downloads8mo agoHugging Face23ai-spatial /GeoSR-Bench GeoSR-Bench Dataset and model weights for the paper: Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration [arXiv] The code is available on GitHub: https://github.com/ai-spatial/GeoSR-Bench Dataset Description GeoSR-Bench directly connects super-resolution (SR) with downstream Earth monitoring tasks, moving beyond conventional fidelity-based evaluation. It comprises spatially co-located… See the full description on the dataset page: https://huggingface.co/datasets/ai-spatial/GeoSR-Bench.textimage-to-image10K<n<100K8 likes2.9k downloads4mo agoHugging Face24OpenDataArena /Spark-234K Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.texttext-generation100K<n<1M60 likes2.5k downloads17d agoHugging Face25argilla /gutenberg_spacy-ner Dataset Card for "gutenberg_spacy-ner" More Information needed textn<1K4 likes2.4k downloads3y agoHugging Face26a8cheng /SpatialRGPT-Benchimage1K<n<10K13 likes2.3k downloads1y agoHugging Face27codesignal /sms-spam-collection SMS Spam Collection v.1 DESCRIPTION The SMS Spam Collection v.1 (hereafter the corpus) is a set of SMS tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged acording being ham (legitimate) or spam. 1.1. Compilation This corpus has been collected from free or free for research sources at the Web: A collection of between 425 SMS spam messages extracted manually from the Grumbletext Web… See the full description on the dataset page: https://huggingface.co/datasets/codesignal/sms-spam-collection.text1K<n<10K1 likes2.1k downloads3y agoHugging Face28RicardoRei /wmt-mqm-error-spans Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context in a form of error spans. Moreover, it contains some hallucinations used in the training of XCOMET models. Please note that this is not an official release of the data and the original data can be found here. The data is organised into 8 columns: src: input text mt: translation ref: reference translation annotations: List… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-error-spans.text100K<n<1M4 likes2k downloads3y agoHugging Face29ekacare /spandan-1M-V1.0-raw Spandan A Large Photoplethysmography (PPG) Signal Dataset of 1 Million+ Indian Subjects In Sanskrit, "Spandan" (स्पन्दन - spandana) represents one of the most fundamental aspects of existence - the rhythmic pulsation that permeates all life. Derived from the root verb "spand" (स्पन्द), meaning "to throb" or "to pulsate," it beautifully captures the essence of the heartbeat. Dataset Overview Spandan is an extensive repository containing over 1 million… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/spandan-1M-V1.0-raw.text100K<n<1M3 likes1.9k downloads2y agoHugging Face30Koushul /spacetravlr SpaceTravLR dataset hub Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports. Layout spacetravlr/ ├── tonsil/ # placeholder / demo gene outputs └── xenium_skin_mixed/ ├── run.toml # shared training config for this cohort ├── manifest.json # sample index and upload metadata ├── sample12/ ├── sample13/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.text10K<n<100K0 likes1.9k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.