datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.thomas-2018-spark-wt
SPARK (wild-type accumulator phenotype): Human-curated and standardized MICs
These data were collated by the authors of:
Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell
Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery
ACS Infectious Diseases 2018 4 (11), 1536-1539
DOI: 10.1021/acsinfecdis.8b00193
We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values,
give succint column… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-wt.polymarket_crypto_derivativesWhole bunch of data from 15 minute crypto markets on Polymarket
amc2023spark2-5-tiny-cpu-repro-v1
Spark2.5 tiny random CPU fixture
Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights.
This is a reproducibility fixture, not a useful language model, distilled model,
quality benchmark, or production registry measurement. No upstream weights, training
data, paid GPU or cloud rentals were used.
Architecture, code and license
Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281.
The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.dgx-spark-benchmarks
DGX Spark LLM Arena benchmarks
Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.safedocs-1M-muse-spark-1.3-first3
SafeDocs first three shards: Muse Spark 1.3
Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b.
PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 8,
"total_frames": 18407,
"total_tasks": 1,
"total_videos": 24,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SparkingGalaxy/so100_test.sparkle_shuffle3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 25050,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_shuffle3.InstaArt-HumanAI
Instagram AI Art vs Human Art: Engagement & Comment Dataset
Dataset Summary
This dataset was created and contributed by Akshaya, Cynthia, Grace, and Soham as part of a project at UC San Diego.
This dataset supports research into how audiences engage with AI-generated art versus human-made art on Instagram, with a specific focus on comment sentiment, reaction types, and engagement patterns. It consists of 40 matched pairs of Instagram posts - one human art post and one… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/InstaArt-HumanAI.spark2-5-tiny-fidelity-root-v1
spark2-5 random CPU fixture root
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/spark2-5-tiny-random-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-fidelity-root-v1.sparkle_project1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 36840,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_project1.sparkle_preview
🤔 About SPARKLE
The work examines how auxiliary information (hints) influences model behavior under RL, studying four types of hints:
Partial Step Scaffolding,
High-level Plans,
External Knowledge, and
Chains of Subproblems.
The SPARKLE dataset provides the structured annotations—plans, knowledge snippets, and decomposed subproblems—that enable this analysis.
These annotations allow researchers to probe how models respond to different forms of auxiliary information and how… See the full description on the dataset page: https://huggingface.co/datasets/sparkle-reasoning/sparkle_preview.kupe-spark-150m-conversationsarcee-ai__Llama-Spark-details
Dataset Card for Evaluation run of arcee-ai/Llama-Spark
Dataset automatically created during the evaluation run of model arcee-ai/Llama-Spark
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Llama-Spark-details.sparkle_shuffle2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 25904,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_shuffle2.spark-capacity-boundary-study
Capacity-boundary optimization with Spark-X2.5-1.7B
This project contains an original evaluation for HER Hack-Astron #6. The report is in DISCUSSION.md. Creating or publishing these artifacts is not an award or payment.
The experiment checks whether increasing a six-item 0/1 knapsack's capacity by one causes the model to find the new optimum. Four seeded item families produce eight mathematical instances, each in English and Chinese. Every prompt runs once with thinking off and… See the full description on the dataset page: https://huggingface.co/datasets/aaaded/spark-capacity-boundary-study.sparkle_project2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 34347,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_project2.dataset_sparkle4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 9085,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/dataset_sparkle4.dataset1_sparkleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 8980,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/dataset1_sparkle.sparkle_project_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 32404,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_project_3.dataset2_sparkleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 23,
"total_frames": 15951,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/dataset2_sparkle.dgx-spark-eval
DGX Spark Model Evaluations
75 Messläufe in fünf Konfigurationen, alle auf einer Maschine gemessen.
Keine Herstellerangaben — jede Zahl stammt aus einem eigenen Lauf. Stand: 2026-08-17.
Die Website zu denselben Daten: https://results.southbyte.de/
Was gemessen wurde
Config
Zeilen
Inhalt
llm_local
20
Sprachmodelle, lokal mit vLLM serviert
llm_saas
28
dieselben Testfälle gegen Frontier-APIs, als Referenzrahmen
guardrails
5
Guard-Modelle gegen einen… See the full description on the dataset page: https://huggingface.co/datasets/SouthByte/dgx-spark-eval.mhpp
MHPP (questions only)
Evaluation-only questions for code/problem-solving. Do not train on test.
Data fields
id (int), question (string), prompt (string), function_name (string), parameters (list[str]), difficulty_types (int)
sparkle_shuffle1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 32,
"total_frames": 26272,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/sparkle_shuffle1.Jacoby746__Proto-Harpy-Spark-v0.1-7B-details
Dataset Card for Evaluation run of Jacoby746/Proto-Harpy-Spark-v0.1-7B
Dataset automatically created during the evaluation run of model Jacoby746/Proto-Harpy-Spark-v0.1-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jacoby746__Proto-Harpy-Spark-v0.1-7B-details.so100-striped-blockdataset_sparkle_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 6578,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/dataset_sparkle_2.thomas-2018-spark-all
SPARK: Human-curated and standardized MICs
These data were collated by the authors of:
Joe Thomas, Marc Navre, Aileen Rubio, and Allan Coukell
Shared Platform for Antibiotic Research and Knowledge: A Collaborative Tool to SPARK Antibiotic Discovery
ACS Infectious Diseases 2018 4 (11), 1536-1539
DOI: 10.1021/acsinfecdis.8b00193
We cleaned the original SPARK dataset to subset the most relevant columns, remove empty values,
give succint column titles, and split by species.
The… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/thomas-2018-spark-all.dataset_sparkle_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 6155,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/itsbyszulaika/dataset_sparkle_1.
