datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPARC-VQA
SPARC VQA
SPARC VQA is the generated spatial VQA training dataset used in the SPARC Qwen3.5 model releases. Each example embeds its image bytes and includes a question, answer, task type, target type, source dataset identifier, split, and JSON metadata.
Raw unfiltered corpus: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw
Ready-to-train split
Use train_filtered_t097_mpo700.parquet for SPARC-only training. This is the processed, release-ready dataset: it… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA.benchmark-noise_ablation
benchmark-noise_ablation
This dataset contains fixed-length noisy speech mixtures generated from LibriSpeech test-clean clean speech and two interference types:
50% DEMAND environmental background noise
50% generated white noise
The dataset is intended for SPARCO noise ablation experiments, including:
AUROC-based SAE noise-related feature selection
binary noise-presence scorer training
scorer threshold calibration
final held-out benchmark evaluation
Sources
Clean… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark-noise_ablation.sparc-rotation-curvesSQL_SparC_Dataset_With_Schema
Dataset Card for "SQL_SparC_Dataset_With_Schema"
More Information needed
benchmark-pitchSPARC-VQA-Raw
SPARC VQA Raw
This repository contains the unfiltered SPARC VQA corpus: 838,211 embedded-image training examples in train.parquet (33.20 GB). Each example contains an image, question, answer, task metadata, source identifier, and annotation metadata including selected_start_score.
Ready-to-train version
For the exact processed SPARC subset used by the released Qwen3.5 models, download train_filtered_t097_mpo700.parquet from irl-kit/SPARC-VQA. It contains 284,909… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw.spar-concepts
SPAR steering concepts
Contrastive texts for building steering directions, used by the SPAR project for the GLP steering evals and for
training and evaluating the cold-diffusion model. A concept is a set of positive texts (target=1, concept
present) and negative texts (target=0, concept absent); its steering direction is the diff-of-means of a model's
activations on the two classes, exactly as in the GLP sentiment eval.
Configs
config
rows
content
texts… See the full description on the dataset page: https://huggingface.co/datasets/basta/spar-concepts.benchmark-sisdrbenchmark-genderbenchmark-speech-f0benchmark-instrumentbenchmark-eventbenchmark-timesparc-galaxiesspar-cot
