CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lkaesberg /SPaRC SPaRC Dataset Website • Solver • Generator A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities. Data Schema Each record (JSON) includes: id (string): unique puzzle identifier difficulty_level (int) & difficulty_score (float) grid_size: { "height": H, "width": W } polyshapes: JSON string mapping shape IDs to binary grids puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112) solution_count (int) and solutions list (with… See the full description on the dataset page: https://huggingface.co/datasets/lkaesberg/SPaRC.tabular1K<n<10K2 likes1.2k downloads2mo agoHugging Face02SPARCO-project /benchmark_DEMAND_noise benchmark_DEMAND_noise This dataset is a segmented subset derived from DEMAND: Diverse Environments Multichannel Acoustic Noise Database. It is prepared for the SPARCO noise ablation benchmark. The intended use is to provide fixed 4-second environmental noise segments for: AUROC-based SAE noise-related feature selection binary noise-presence scorer training scorer threshold calibration final held-out benchmark evaluation Source Original source: DEMAND: Diverse… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark_DEMAND_noise.audioaudio-classification1K<n<10K0 likes354 downloads4mo agoHugging Face03irl-kit /SPARC-VQA SPARC VQA SPARC VQA is the generated spatial VQA training dataset used in the SPARC Qwen3.5 model releases. Each example embeds its image bytes and includes a question, answer, task type, target type, source dataset identifier, split, and JSON metadata. Raw unfiltered corpus: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw Ready-to-train split Use train_filtered_t097_mpo700.parquet for SPARC-only training. This is the processed, release-ready dataset: it… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA.image100K<n<1M0 likes134 downloads1mo agoHugging Face04aherntech /sparc Dataset Card for SParC SParC is a context-dependant multi-turn version of the Spider task 1.0. This dataset provides a chat-bot oriented test set for text-to-sql problems. Additional details may be obtained in the paper: https://arxiv.org/abs/1906.02285 Paper Abstract We present SParC, a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences (12k+ individual questions annotated with SQL queries). It is obtained from… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/sparc.text1K<n<10K1 likes128 downloads3y agoHugging Face05average-developer /stocks-SPARC-1D-candlesn<1K0 likes128 downloads19h agoHugging Face06jamilalthani1 /SPARC SPaRC Dataset Website • Solver • Generator A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities. Data Schema Each record (JSON) includes: id (string): unique puzzle identifier difficulty_level (int) & difficulty_score (float) grid_size: { "height": H, "width": W } polyshapes: JSON string mapping shape IDs to binary grids puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112) solution_count (int) and solutions list (with index… See the full description on the dataset page: https://huggingface.co/datasets/jamilalthani1/SPARC.tabular1K<n<10K0 likes105 downloads10mo agoHugging Face07SPARCO-project /benchmark-noise_ablation benchmark-noise_ablation This dataset contains fixed-length noisy speech mixtures generated from LibriSpeech test-clean clean speech and two interference types: 50% DEMAND environmental background noise 50% generated white noise The dataset is intended for SPARCO noise ablation experiments, including: AUROC-based SAE noise-related feature selection binary noise-presence scorer training scorer threshold calibration final held-out benchmark evaluation Sources Clean… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark-noise_ablation.audioaudio-to-audio1K<n<10K0 likes83 downloads4mo agoHugging Face08jellyChiru /SParC@InProceedings{Yu&al.19, title = {SParC: Cross-Domain Semantic Parsing in Context}, author = {Tao Yu and Rui Zhang and Michihiro Yasunaga and Yi Chern Tan and Xi Victoria Lin and Suyi Li and Heyang Er, Irene Li and Bo Pang and Tao Chen and Emily Ji and Shreya Dixit and David Proctor and Sungrok Shim and Jonathan Kraft, Vincent Zhang and Caiming Xiong and Richard Socher and Dragomir Radev}, booktitle = {Proceedings of the 57th Annual Meeting of the Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/jellyChiru/SParC.text1K<n<10K2 likes66 downloads3y agoHugging Face09scidata-hub /sparc-rotation-curvestabular1K<n<10K0 likes66 downloads3mo agoHugging Face10AayushShah /SQL_SparC_Dataset_With_Schema Dataset Card for "SQL_SparC_Dataset_With_Schema" More Information needed text1K<n<10K3 likes64 downloads3y agoHugging Face11SPARCO-project /benchmark-pitchaudio1K<n<10K0 likes64 downloads4mo agoHugging Face12irl-kit /SPARC-VQA-Raw SPARC VQA Raw This repository contains the unfiltered SPARC VQA corpus: 838,211 embedded-image training examples in train.parquet (33.20 GB). Each example contains an image, question, answer, task metadata, source identifier, and annotation metadata including selected_start_score. Ready-to-train version For the exact processed SPARC subset used by the released Qwen3.5 models, download train_filtered_t097_mpo700.parquet from irl-kit/SPARC-VQA. It contains 284,909… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw.image100K<n<1M0 likes55 downloads1mo agoHugging Face13irl-kit /sparc-droid-annotations SPARC DROID annotations SPARC annotations for DROID. This repository contains annotations only. It does not redistribute DROID images, videos, actions, or robot states; obtain the source dataset separately. Each JSONL row describes one interaction subtask from one camera view. Dense arrays live in HDF5 sidecars referenced by that row. All coordinates, masks, and frame indices refer to the original uncropped, unresized DROID camera frames. Release contents… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/sparc-droid-annotations.object-detection0 likes54 downloads26d agoHugging Face14avandar2024 /sparc0 likes52 downloads1mo agoHugging Face15basta /spar-concepts SPAR steering concepts Contrastive texts for building steering directions, used by the SPAR project for the GLP steering evals and for training and evaluating the cold-diffusion model. A concept is a set of positive texts (target=1, concept present) and negative texts (target=0, concept absent); its steering direction is the diff-of-means of a model's activations on the two classes, exactly as in the GLP sentiment eval. Configs config rows content texts… See the full description on the dataset page: https://huggingface.co/datasets/basta/spar-concepts.text10K<n<100K1 likes47 downloads3d agoHugging Face16SPARCO-project /benchmark-sisdraudio1K<n<10K0 likes43 downloads5mo agoHugging Face17SPARCO-project /benchmark-genderaudio1K<n<10K0 likes42 downloads4mo agoHugging Face18SPARCO-project /benchmark-speech-f0audio1K<n<10K0 likes34 downloads2mo agoHugging Face19SPARCO-project /benchmark-instrumentaudio1K<n<10K0 likes27 downloads4mo agoHugging Face20adarshsahu27 /sparc-llama2text1K<n<10K0 likes11 downloads2y agoHugging Face21SPARCO-project /benchmark-eventaudion<1K0 likes10 downloads5mo agoHugging Face22SPARCO-project /benchmark-timeaudio1K<n<10K0 likes9 downloads4mo agoHugging Face23scidata-hub /sparc-galaxiestabularn<1K0 likes9 downloads3mo agoHugging Face24jiyatai /spar-cotimage10K<n<100K0 likes7 downloads4mo agoHugging Face25Romila2036 /sparc_project0 likes5 downloads1mo agoHugging Face26znyang /sparcimagequestion-answering1K<n<10K1 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.