CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes34k downloads1y agoHugging Face02albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads6d agoHugging Face03stdKonjac /Sparkle Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.imagetext-to-video100K<n<1M1 likes4.4k downloads5mo agoHugging Face04OpenDataArena /Spark-234K Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.texttext-generation100K<n<1M72 likes3k downloads19d agoHugging Face05scbirlab /thomas-2018-spark-wt SPARK (wild-type accumulator phenotype): Human-curated and standardized MICs These data were collated by the authors of: Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery ACS Infectious Diseases 2018 4 (11), 1536-1539 DOI: 10.1021/acsinfecdis.8b00193 We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values, give succint column… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-wt.tabulartext-classification100K<n<1M0 likes1k downloads11mo agoHugging Face06LGB666 /SageLM-CosyVoice2-SparkTTStext10K<n<100K0 likes509 downloads10mo agoHugging Face07LGB666 /SageLM-Spark-TTStext10K<n<100K0 likes423 downloads10mo agoHugging Face08sparklabutah /TimeWarp-GPT5-Tracesimage0 likes367 downloads7mo agoHugging Face09stdKonjac /Sparkle-Bench Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle-Bench.imagetext-to-videon<1K1 likes285 downloads5mo agoHugging Face10sparks-solutions /AutomotiveUI-Bench-4K AutomotiveUI-Bench-4K Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems. Key Features: Serves as a validation benchmark for automotive UI. Scope: Covers 15 automotive brands/OEMs, model years 2018-2025. Image Source: Primarily photographs of IVI displays (due to screenshot limitations in most vehicles), with some direct screenshots (e.g., Android Auto). Annotation Classes: Test Action: Bounding box + imperative… See the full description on the dataset page: https://huggingface.co/datasets/sparks-solutions/AutomotiveUI-Bench-4K.imagevisual-question-answering1K<n<10K6 likes222 downloads1y agoHugging Face11gittensor-model-hub /sparkproof-miningtext1K<n<10K0 likes204 downloads2mo agoHugging Face12sparkle-reasoning /amc2023tabularn<1K0 likes173 downloads1y agoHugging Face13malaiwah /spark2-5-tiny-cpu-repro-v1 Spark2.5 tiny random CPU fixture Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights. This is a reproducibility fixture, not a useful language model, distilled model, quality benchmark, or production registry measurement. No upstream weights, training data, paid GPU or cloud rentals were used. Architecture, code and license Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281. The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.tabularn<1K0 likes147 downloads18d agoHugging Face14Djangodevreng /dgx-spark-benchmarks DGX Spark LLM Arena benchmarks Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.tabularn<1K1 likes131 downloads7d agoHugging Face15albertklorer /safedocs-1M-muse-spark-1.3-first3 SafeDocs first three shards: Muse Spark 1.3 Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b. PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.tabularn<1K0 likes125 downloads9d agoHugging Face16sparklessszzz /NewsLensSync Dataset Card for NewsLensSync Dataset Description This dataset, named NewsLensSync, contains a curated collection of news articles, sourced from trusted domains such as BBC, Reuters, AP News, NPR, PBS, The Guardian, WSJ, NY Times, and ProPublica. Each article includes both the original content and a synthetic "falsified" version of the article description, generated using a transformer-based negation model. The dataset is designed for research in misinformation… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/NewsLensSync.texttext-classification10K<n<100K2 likes110 downloads25d agoHugging Face17michsethowusu /spark-tts-twi-preppedtext10K<n<100K0 likes80 downloads4mo agoHugging Face18sparklessszzz /InstaArt-HumanAI Instagram AI Art vs Human Art: Engagement & Comment Dataset Dataset Summary This dataset was created and contributed by Akshaya, Cynthia, Grace, and Soham as part of a project at UC San Diego. This dataset supports research into how audiences engage with AI-generated art versus human-made art on Instagram, with a specific focus on comment sentiment, reaction types, and engagement patterns. It consists of 40 matched pairs of Instagram posts - one human art post and one… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/InstaArt-HumanAI.tabulartext-classificationn<1K1 likes79 downloads7mo agoHugging Face19malaiwah /spark2-5-tiny-fidelity-root-v1 spark2-5 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/spark2-5-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-fidelity-root-v1.tabularn<1K0 likes79 downloads18d agoHugging Face20SparkSupernova /nova-industry-benchmark Nova Industry Benchmark 22 curated questions across six categories, used to evaluate NovaLiveSystem. This dataset contains questions only. Model outputs live in SparkSupernova/nova-industry-benchmark-results, joined on id. Why the split Run outputs used to be stored here as additional splits. Each model version wrote a different set of columns, so adding the v5 run left this dataset with two splits whose schemas disagreed, and load_dataset raised… See the full description on the dataset page: https://huggingface.co/datasets/SparkSupernova/nova-industry-benchmark.texttext-generationn<1K0 likes75 downloads2mo agoHugging Face21SparkleDark /Everything_about_dogs Dataset Card for Everything About Dogs Dataset This dataset is built from the book Everything About Dogs Everything about dogs is a book by AL G Eberhart. The book contains topics on diseases, how to feed, environment to provide, how to take care of pups and everything about dogs. Motivation In Memory of Caspu Dogs can't speak but they can express their pain. However this expression is often misunderstood or ignored. My I aim is to build a retrieval… See the full description on the dataset page: https://huggingface.co/datasets/SparkleDark/Everything_about_dogs.texttext-generation10K<n<100K2 likes72 downloads1y agoHugging Face22Sparkteknologiiii /Lab2text1M<n<10M1 likes71 downloads4d agoHugging Face23Dude311 /spark-math-audit-20260911 Spark-X2.5: solving and auditing misleading worked solutions Status: experiment running; not a completed competition entry yet. Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.texttext-generationn<1K0 likes69 downloads14d agoHugging Face24G3nadh /dgx-spark-benchmarks DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1 8B 4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.texttext-generationn<1K1 likes67 downloads6mo agoHugging Face25topyun /SPARK ⚡ SPARK (multi-vision Sensor Perception And Reasoning benchmarK) 🌐 github | 🤗 Dataset | 📃 Paper Dataset Details SPARK can reduce the fundamental multi-vision sensor information gap between images and multi-vision sensors. We generated 6,248 vision-language test samples automatically to investigate multi-vision sensory perception and multi-vision sensory reasoning on physical sensor knowledge proficiency across different formats, covering different types of… See the full description on the dataset page: https://huggingface.co/datasets/topyun/SPARK.image1K<n<10K15 likes65 downloads2y agoHugging Face26sparkle-reasoning /sparkle_preview 🤔 About SPARKLE The work examines how auxiliary information (hints) influences model behavior under RL, studying four types of hints: Partial Step Scaffolding, High-level Plans, External Knowledge, and Chains of Subproblems. The SPARKLE dataset provides the structured annotations—plans, knowledge snippets, and decomposed subproblems—that enable this analysis. These annotations allow researchers to probe how models respond to different forms of auxiliary information and how… See the full description on the dataset page: https://huggingface.co/datasets/sparkle-reasoning/sparkle_preview.tabular10K<n<100K1 likes62 downloads10mo agoHugging Face27pocharlies /dgx-spark-moe-benchmarks Four MoE models on a DGX Spark: speed, tool-calling, and what actually breaks Full measurement campaign on NVIDIA DGX Spark (GB10, 128 GB unified, ~273 GB/s), vLLM 0.23.1rc1.dev301+g04c2a8dea, arm64/sm121. Every number here is measured on this hardware, with the raw evidence included. The headline: on synthetic tool-calling benchmarks all four models score 91-95 %. In a real coding agent, three of them score 0-1 out of 14 and one scores 11 out of 14. If you pick a model from the… See the full description on the dataset page: https://huggingface.co/datasets/pocharlies/dgx-spark-moe-benchmarks.textn<1K0 likes58 downloads2mo agoHugging Face28eruantion87 /craft-arithmetic24-spark-budget-study CraftArithmetic24 and Spark CPU budget study Twenty-four original synthetic arithmetic questions across six related problem families. Each family contains a base problem, paraphrase, changed-number problem and irrelevant-detail variant. See dataset.jsonl and dataset_metadata.json for prompts, exact labels, version and hash. The original questions and labels are released under CC0-1.0. The original evaluation code is released under MIT. Model-generated outputs are included as… See the full description on the dataset page: https://huggingface.co/datasets/eruantion87/craft-arithmetic24-spark-budget-study.textquestion-answeringn<1K0 likes58 downloads12d agoHugging Face29anuj-inavlabs /kupe-spark-150m-conversationstabular1K<n<10K0 likes51 downloads6d agoHugging Face30SALT-NLP /SparkMe-SyntheticUsers SparkMe-SyntheticUsers Synthetic user profiles for evaluating AI interview systems, released alongside the SparkMe. Each profile represents a simulated workforce participant with demographic metadata, a shuffled list of persona facts, and structured ground-truth interview notes across 10 topics covering the impact of AI in the workplace. Dataset Description The 200 profiles were generated from WorkBank worker seed data using SparkMe's user agent pipeline. Each user has:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SparkMe-SyntheticUsers.texttext-generationn<1K1 likes48 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.