datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxbox
VoxBox
This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion.
Dataset Structure
.
├── audios/
│ └── aishell-3/ # Audio files (organised by sub-corpus)
│ └── ...
└── metadata/
├── aishell-3.jsonl
├── casia.jsonl
├── commonvoice_cn.jsonl
├── ...
└── wenetspeech4tts.jsonl # JSONL metadata files
Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.Sparkle
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.Spark-234K
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.CosyVoice2-SparkTTSthomas-2018-spark-wt
SPARK (wild-type accumulator phenotype): Human-curated and standardized MICs
These data were collated by the authors of:
Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell
Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery
ACS Infectious Diseases 2018 4 (11), 1536-1539
DOI: 10.1021/acsinfecdis.8b00193
We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values,
give succint column… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-wt.GPT-5.3-Codex-Spark-CodexThis dataset was generated using teich by TeichAI
GPT-5.3-Codex-Spark Codex
This directory contains raw agent trace files generated by teich.
JSONL files: 200
Model metadata: gpt-5.3-codex-spark
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.3-Codex-Spark-Codex.kupe-spark-asr-270m-data
kupe-spark-asr-270m — data
Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec).
Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi)
Configs
audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet).
mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training.
Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.EditHF-1M
Code Resources
For evaluation code and model (EditHF and EditHF-Reward), please visit our GitHub repository: GitHub Repository
Data Structure
EditHF-1M/
├── source/
│ ├── category1/
│ │ ├── 1.jpg
│ │ ├── 2.jpg
│ │ └── ...
│ ├── category2/
│ └── ...
│
├── edited/
│ ├── model1/
│ │ ├── category1/
│ │ │ ├── 1.jpg
│ │ │ ├── 2.jpg
│ │ │ └── ...
│ │ ├── category2/
│ │ └── ...
│ ├── model2/
│ └── ...
│
├── scores/
│… See the full description on the dataset page: https://huggingface.co/datasets/sparkling621/EditHF-1M.timewarp-env-data
TimeWarp Environment Data
Data files consumed by sparklabutah/timewarp setup.sh when provisioning the Wiki, News, and Shop environments.
Contents
File
Used by
wiki_index.pkl
env/wiki
news_index.pkl
env/news
webshop/items_shuffle_1000.json
env/webshop (setup.sh -d small) — product info, 1k subset
webshop/items_ins_v2_1000.json
env/webshop (setup.sh -d small) — product attributes, 1k subset
webshop/items_shuffle.json
env/webshop (setup.sh -d all)… See the full description on the dataset page: https://huggingface.co/datasets/sparklabutah/timewarp-env-data.SageLM-CosyVoice2-SparkTTSSageLM-Spark-TTSRoboFactory_asset
RoboFactory Asset Dataset Card
This repository contains basic assets for RoboFactory.
SPARK-2022
SPARK 2022 — Stream 1 (Spacecraft Detection)
Stream 1 of the SPARK 2022 dataset (SPAcecraft Recognition leveraging Knowledge
of the space environment): space-borne imagery of 10 spacecraft plus a debris
class, for object detection and classification. Each image contains exactly one
target annotated with a single bounding box and class label.
Dataset summary
Images
110,000 JPEG, 1024 × 1024, RGB
Annotations
1 bounding box + class per image
Classes… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2022.spark-plug-anomaly-detection
Spark Plug Anomaly Detection
Images of spark plugs for use in visual anomaly detection with Edge Impulse’s FOMO-AD learning block. The model is trained only on normal samples and flags any deviation as an anomaly.
Input: Images (96x96)
Classes: normal (train/test), anomaly (test only)
Use case: Embedded anomaly detection in predictive maintenance
Model trained and demonstrated on Edge Impulse.
Looking for multi-class condition labels?See the companion dataset: Spark Plug… See the full description on the dataset page: https://huggingface.co/datasets/eoinedge/spark-plug-anomaly-detection.TimeWarp-GPT5-TracesSPARK_PDI_Trajectory
SPARK PDI Trajectory
Trajectory-level artifacts released alongside the paper
Evidence Over Plans: Online Trajectory Verification for Skill Distillation.
This dataset contains the raw execution trajectories, exploration memos, and distilled SKILL.md
documents produced by the SPARK skill-generation pipeline. It is the primary data source used to compute the
Posterior Distillation Index (PDI) — a trajectory-level score that measures whether a distilled skill is
grounded in posterior… See the full description on the dataset page: https://huggingface.co/datasets/EtaYang10th/SPARK_PDI_Trajectory.Sparkle-Bench
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle-Bench.sparkproof-miningpolymarket_crypto_derivativesWhole bunch of data from 15 minute crypto markets on Polymarket
AutomotiveUI-Bench-4K
AutomotiveUI-Bench-4K
Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems.
Key Features:
Serves as a validation benchmark for automotive UI.
Scope: Covers 15 automotive brands/OEMs, model years 2018-2025.
Image Source: Primarily photographs of IVI displays (due to screenshot limitations in most vehicles), with some direct screenshots (e.g., Android Auto).
Annotation Classes:
Test Action: Bounding box + imperative… See the full description on the dataset page: https://huggingface.co/datasets/sparks-solutions/AutomotiveUI-Bench-4K.amc2023Spark-AnomalyGen-USD
Dataset Overview
Dataset Description:
The asset in question is the USD along with the components. Full PCBA scene (spark_lighting.usd) with an authored AOI ring-light rig (aoi_ring_light.usda) and camera — ready for synthetic data generation rendering. USD is derived from the underlying CAD design.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
05/30/2026
Version:
1.0
License/Terms of Use:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Spark-AnomalyGen-USD.SPARK-2026
Dataset Card for SPARK 2026/2024
Although SPARK 2024 Stream-1 and SPARK 2026 Stream-1 utilize the exact same dataset, the underlying tasks differ. While SPARK 2024 focused exclusively on spacecraft component semantic segmentation, SPARK 2026 expands the objective to multi-task learning paired with efficient model architecture design.
SPARK 2026 is a dataset of spacecraft imagery with bounding box labels and segmentation masks, released as part of the SPARK 2026 Challenge for… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2026.DVF
Dataset Card for DVF
This is the dataset of Diffusion Video Forensics (DVF) from On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
. The github link is here.
Dataset Details
This dataset provides the reconstructed frames of the full DVF, as well as the cached MM Representations.
SPARK-2021
SPARK-2021: SPAcecraft Recognition leveraging Knowledge of space environment
SPARK is a large-scale multi-modal (RGB + depth) synthetic image dataset for space object
recognition and detection, generated under a photo-realistic space simulation environment.
It was released by the CVI² group at SnT, University of Luxembourg
in the context of the SPARK Challenge at IEEE ICIP 2021.
The dataset targets Space Situational Awareness (SSA) applications — on-orbit servicing,
active… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPARK-2021.tmp-Infinity-Instruct-3M-sparkdgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.DVF_PPSPARK-2025
