CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes45k downloads1y agoHugging Face02albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads2d agoHugging Face03stdKonjac /Sparkle Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.imagetext-to-video100K<n<1M1 likes4.4k downloads5mo agoHugging Face04OpenDataArena /Spark-234K Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.texttext-generation100K<n<1M55 likes2.3k downloads15d agoHugging Face05LGB666 /CosyVoice2-SparkTTSaudion<1K0 likes1.5k downloads1y agoHugging Face06scbirlab /thomas-2018-spark-wt SPARK (wild-type accumulator phenotype): Human-curated and standardized MICs These data were collated by the authors of: Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery ACS Infectious Diseases 2018 4 (11), 1536-1539 DOI: 10.1021/acsinfecdis.8b00193 We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values, give succint column… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-wt.tabulartext-classification100K<n<1M0 likes1.2k downloads11mo agoHugging Face07AletheiaResearch /GPT-5.3-Codex-Spark-CodexThis dataset was generated using teich by TeichAI GPT-5.3-Codex-Spark Codex This directory contains raw agent trace files generated by teich. JSONL files: 200 Model metadata: gpt-5.3-codex-spark Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.3-Codex-Spark-Codex.text-generation10K<n<100K4 likes882 downloads3mo agoHugging Face08anuj-inavlabs /kupe-spark-asr-270m-data kupe-spark-asr-270m — data Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec). Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi) Configs audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet). mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training. Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.audio0 likes663 downloads17d agoHugging Face09sparkling621 /EditHF-1M Code Resources For evaluation code and model (EditHF and EditHF-Reward), please visit our GitHub repository: GitHub Repository Data Structure EditHF-1M/ ├── source/ │ ├── category1/ │ │ ├── 1.jpg │ │ ├── 2.jpg │ │ └── ... │ ├── category2/ │ └── ... │ ├── edited/ │ ├── model1/ │ │ ├── category1/ │ │ │ ├── 1.jpg │ │ │ ├── 2.jpg │ │ │ └── ... │ │ ├── category2/ │ │ └── ... │ ├── model2/ │ └── ... │ ├── scores/ │… See the full description on the dataset page: https://huggingface.co/datasets/sparkling621/EditHF-1M.1 likes580 downloads2mo agoHugging Face10sparklabutah /timewarp-env-data TimeWarp Environment Data Data files consumed by sparklabutah/timewarp setup.sh when provisioning the Wiki, News, and Shop environments. Contents File Used by wiki_index.pkl env/wiki news_index.pkl env/news webshop/items_shuffle_1000.json env/webshop (setup.sh -d small) — product info, 1k subset webshop/items_ins_v2_1000.json env/webshop (setup.sh -d small) — product attributes, 1k subset webshop/items_shuffle.json env/webshop (setup.sh -d all)… See the full description on the dataset page: https://huggingface.co/datasets/sparklabutah/timewarp-env-data.0 likes489 downloads2mo agoHugging Face11LGB666 /SageLM-CosyVoice2-SparkTTStext10K<n<100K0 likes475 downloads10mo agoHugging Face12LGB666 /SageLM-Spark-TTStext10K<n<100K0 likes412 downloads10mo agoHugging Face13sparklexfantasy /RoboFactory_asset RoboFactory Asset Dataset Card This repository contains basic assets for RoboFactory. image1K<n<10K0 likes379 downloads1y agoHugging Face14CVI2-UniLU /SPARK-2022 SPARK 2022 — Stream 1 (Spacecraft Detection) Stream 1 of the SPARK 2022 dataset (SPAcecraft Recognition leveraging Knowledge of the space environment): space-borne imagery of 10 spacecraft plus a debris class, for object detection and classification. Each image contains exactly one target annotated with a single bounding box and class label. Dataset summary Images 110,000 JPEG, 1024 × 1024, RGB Annotations 1 bounding box + class per image Classes… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2022.imageobject-detection100K<n<1M2 likes347 downloads2mo agoHugging Face15eoinedge /spark-plug-anomaly-detection Spark Plug Anomaly Detection Images of spark plugs for use in visual anomaly detection with Edge Impulse’s FOMO-AD learning block. The model is trained only on normal samples and flags any deviation as an anomaly. Input: Images (96x96) Classes: normal (train/test), anomaly (test only) Use case: Embedded anomaly detection in predictive maintenance Model trained and demonstrated on Edge Impulse. Looking for multi-class condition labels?See the companion dataset: Spark Plug… See the full description on the dataset page: https://huggingface.co/datasets/eoinedge/spark-plug-anomaly-detection.imageimage-classification1K<n<10K1 likes333 downloads3mo agoHugging Face16sparklabutah /TimeWarp-GPT5-Tracesimage0 likes330 downloads7mo agoHugging Face17EtaYang10th /SPARK_PDI_Trajectory SPARK PDI Trajectory Trajectory-level artifacts released alongside the paper Evidence Over Plans: Online Trajectory Verification for Skill Distillation. This dataset contains the raw execution trajectories, exploration memos, and distilled SKILL.md documents produced by the SPARK skill-generation pipeline. It is the primary data source used to compute the Posterior Distillation Index (PDI) — a trajectory-level score that measures whether a distilled skill is grounded in posterior… See the full description on the dataset page: https://huggingface.co/datasets/EtaYang10th/SPARK_PDI_Trajectory.text-generationn<1K2 likes262 downloads4mo agoHugging Face18stdKonjac /Sparkle-Bench Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle-Bench.imagetext-to-videon<1K1 likes255 downloads5mo agoHugging Face19gittensor-model-hub /sparkproof-miningtext1K<n<10K0 likes253 downloads2mo agoHugging Face20Sparkoot /polymarket_crypto_derivativesWhole bunch of data from 15 minute crypto markets on Polymarket tabular100M<n<1B0 likes237 downloads8mo agoHugging Face21sparks-solutions /AutomotiveUI-Bench-4K AutomotiveUI-Bench-4K Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems. Key Features: Serves as a validation benchmark for automotive UI. Scope: Covers 15 automotive brands/OEMs, model years 2018-2025. Image Source: Primarily photographs of IVI displays (due to screenshot limitations in most vehicles), with some direct screenshots (e.g., Android Auto). Annotation Classes: Test Action: Bounding box + imperative… See the full description on the dataset page: https://huggingface.co/datasets/sparks-solutions/AutomotiveUI-Bench-4K.imagevisual-question-answering1K<n<10K6 likes229 downloads1y agoHugging Face22sparkle-reasoning /amc2023tabularn<1K0 likes210 downloads1y agoHugging Face23nvidia /Spark-AnomalyGen-USD Dataset Overview Dataset Description: The asset in question is the USD along with the components. Full PCBA scene (spark_lighting.usd) with an authored AOI ring-light rig (aoi_ring_light.usda) and camera — ready for synthetic data generation rendering. USD is derived from the underlying CAD design. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: 05/30/2026 Version: 1.0 License/Terms of Use:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Spark-AnomalyGen-USD.imagen<1K4 likes209 downloads4mo agoHugging Face24CVI2-UniLU /SPARK-2026 Dataset Card for SPARK 2026/2024 Although SPARK 2024 Stream-1 and SPARK 2026 Stream-1 utilize the exact same dataset, the underlying tasks differ. While SPARK 2024 focused exclusively on spacecraft component semantic segmentation, SPARK 2026 expands the objective to multi-task learning paired with efficient model architecture design. SPARK 2026 is a dataset of spacecraft imagery with bounding box labels and segmentation masks, released as part of the SPARK 2026 Challenge for… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2026.imageobject-detection100K<n<1M1 likes177 downloads2mo agoHugging Face25sparklexfantasy /DVF Dataset Card for DVF This is the dataset of Diffusion Video Forensics (DVF) from On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection . The github link is here. Dataset Details This dataset provides the reconstructed frames of the full DVF, as well as the cached MM Representations. 0 likes170 downloads8mo agoHugging Face26CVI2-UniLU /SPARK-2021 SPARK-2021: SPAcecraft Recognition leveraging Knowledge of space environment SPARK is a large-scale multi-modal (RGB + depth) synthetic image dataset for space object recognition and detection, generated under a photo-realistic space simulation environment. It was released by the CVI² group at SnT, University of Luxembourg in the context of the SPARK Challenge at IEEE ICIP 2021. The dataset targets Space Situational Awareness (SSA) applications — on-orbit servicing, active… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2021.image-classification100K<n<1M2 likes168 downloads2mo agoHugging Face27lhoestq /tmp-Infinity-Instruct-3M-spark0 likes167 downloads2y agoHugging Face28Djangodevreng /dgx-spark-benchmarks DGX Spark LLM Arena benchmarks Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.tabularn<1K1 likes162 downloads4d agoHugging Face29sparklexfantasy /DVF_PP0 likes155 downloads11mo agoHugging Face30introvoyz041 /SPARK-20250 likes155 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.