sparc
Datasets
All datasets matching “sparc”SPaRC
SPaRC Dataset
Website • Solver • Generator
A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities.
Data Schema
Each record (JSON) includes:
id (string): unique puzzle identifier
difficulty_level (int) & difficulty_score (float)
grid_size: { "height": H, "width": W }
polyshapes: JSON string mapping shape IDs to binary grids
puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112)
solution_count (int) and solutions list (with… See the full description on the dataset page: https://huggingface.co/datasets/lkaesberg/SPaRC.benchmark_DEMAND_noise
benchmark_DEMAND_noise
This dataset is a segmented subset derived from DEMAND: Diverse Environments Multichannel Acoustic Noise Database.
It is prepared for the SPARCO noise ablation benchmark. The intended use is to provide fixed 4-second environmental noise segments for:
AUROC-based SAE noise-related feature selection
binary noise-presence scorer training
scorer threshold calibration
final held-out benchmark evaluation
Source
Original source:
DEMAND: Diverse… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark_DEMAND_noise.SPARC-VQA
SPARC VQA
SPARC VQA is the generated spatial VQA training dataset used in the SPARC Qwen3.5 model releases. Each example embeds its image bytes and includes a question, answer, task type, target type, source dataset identifier, split, and JSON metadata.
Raw unfiltered corpus: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw
Ready-to-train split
Use train_filtered_t097_mpo700.parquet for SPARC-only training. This is the processed, release-ready dataset: it… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA.sparc
Dataset Card for SParC
SParC is a context-dependant multi-turn version of the Spider task 1.0.
This dataset provides a chat-bot oriented test set for text-to-sql problems. Additional details may be obtained in the paper:
https://arxiv.org/abs/1906.02285
Paper Abstract
We present SParC, a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences (12k+ individual questions annotated with SQL queries). It is obtained from… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/sparc.stocks-SPARC-1D-candlesSPARC
SPaRC Dataset
Website • Solver • Generator
A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities.
Data Schema
Each record (JSON) includes:
id (string): unique puzzle identifier
difficulty_level (int) & difficulty_score (float)
grid_size: { "height": H, "width": W }
polyshapes: JSON string mapping shape IDs to binary grids
puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112)
solution_count (int) and solutions list (with index… See the full description on the dataset page: https://huggingface.co/datasets/jamilalthani1/SPARC.
