CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenDriveLab /SparseVideoNav SparseVideoNav Datasets This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav: BVN: Beyond-the-View Navigation. IFN: Instruction-Following Navigation. Project links: Project page: https://opendrivelab.com/SparseVideoNav GitHub: https://github.com/OpenDriveLab/SparseVideoNav Paper: https://arxiv.org/abs/2602.05827 Dataset Summary SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.tabularrobotics10K<n<100K4 likes6.1k downloads27d agoHugging Face02SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes6k downloads5y agoHugging Face03lkaesberg /SPaRC SPaRC Dataset Website • Solver • Generator A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities. Data Schema Each record (JSON) includes: id (string): unique puzzle identifier difficulty_level (int) & difficulty_score (float) grid_size: { "height": H, "width": W } polyshapes: JSON string mapping shape IDs to binary grids puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112) solution_count (int) and solutions list (with… See the full description on the dataset page: https://huggingface.co/datasets/lkaesberg/SPaRC.tabular1K<n<10K2 likes1k downloads2mo agoHugging Face04SivanSX /spatialtesttabular100K<n<1M0 likes549 downloads2y agoHugging Face05spade-rl /SPADE-Environment-Pool-GPT5.5-ToolUse SPARE GPT-5.5 Multi-Turn Tool-Use Games v1 A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training. Training alignment Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration 400 rollouts x 24 games/rollout = 9,600 no-reuse games required 11,039 validated games provide 1,439 games of headroom Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.tabularreinforcement-learning10K<n<100K1 likes542 downloads29d agoHugging Face06SZLHOLDINGS /uds-spans-receipts Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. UDS Spans Receipts — OTel Governance Audit Log Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI. Append-only audit log of DSSE-signed OpenTelemetry spans emitted by the UDS mesh governance layer. Each span record includes: operation… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-spans-receipts.tabularothern<1K0 likes540 downloads23d agoHugging Face07tomhodemon /grounded-visual-spatial-reasoning Grounded Visual Spatial Reasoning Code for generating the annotations can be found here: github.com Dataset Summary This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image. Data instance Each sample instance has the following structure: Field Type Description image_file string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.image10K<n<100K2 likes354 downloads1y agoHugging Face08nickh007 /sparam-conformance sparam-conformance 📖 Documentation site — the portfolio narrative, the concepts, a full walkthrough, and what all of this proves (and does not). A labelled corpus of S-parameter networks with ground-truth physical verdicts — and a scorer that grades any checker against it. Why this exists There is no public dataset of physically invalid S-parameter files. Everyone building an RF validation tool tests it on files that happen to be lying around, which means… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/sparam-conformance.tabulartabular-classificationn<1K0 likes288 downloads1mo agoHugging Face09maeyounes /SparseCraft-dataset SparseCraft [ECCV'24] SparseCraft: Few-Shot Neural Reconstruction through Stereopsis Guided Geometric Linearization Project DTU Dataset We provide preprocessed DTU data and results for the tasks of novel view synthesis and surface reconstruction. It contains the following directories: sparsecraft_data ├── nvs # Novel View Synthesis task data and results │ └── mvs_data │ ├── scan103 │ ├── ... │ └── results # Results for training using… See the full description on the dataset page: https://huggingface.co/datasets/maeyounes/SparseCraft-dataset.imagen<1K0 likes255 downloads2y agoHugging Face10suitai /salabs-robotics-spatial-topology-v9 🤖 SALabs 768-D Continuous Lie SE(3) Robotics & Spatial Manifold Topology Dataset (v9.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($1,500 USD) & Instant 391.56MB Master DownloadInstant download of the full 391.56MB Enterprise JSONL matrix containing 50,000+ continuous Lie $SE(3)$ manifold trajectories, singularity-free Bishop Frame metrics, and commercial license certificate. 🌟 Executive Summary The SALabs Robotics & Spatial… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-robotics-spatial-topology-v9.tabularrobotics1K<n<10K1 likes209 downloads17d agoHugging Face11placingholocaust /spacy-project 📚 Placing the Holocaust Weasel (spacy) Project This is the official spaCy project for the Placing the Holocaust Project. This project houses our data and our Python scripts for converting data, serializing it, training 4 different spaCy models with it, and evaluating those models. It also contains all the metrics from v. 0.0.1. For this project, we are using spaCy v. 3.7.4. Project Overview Studying experiences of the Holocaust should not be limited to what happened in… See the full description on the dataset page: https://huggingface.co/datasets/placingholocaust/spacy-project.tabular10K<n<100K0 likes172 downloads2y agoHugging Face12suitai /salabs-virtual-spatial-digitaltwin-v8 🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate. 🌟 Executive Summary The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.tabularother1K<n<10K1 likes136 downloads17d agoHugging Face13SpatiaOS /P3D-Bench P3D-Bench P3D-Bench is the lightweight data release for P3D-Bench, a benchmark for executable parametric 3D generation. It follows the three task splits used in the paper: Text-to-3D: 400 Text2CAD-derived single-part cases. Image-to-3D: 400 Fusion 360 Gallery assembly cases. Assembly-3D: 203 verified Fusion 360 Gallery cases with assembly-level and part-level annotations. This repository publishes the final benchmark UID lists, the P3D-derived text/assembly annotations, the… See the full description on the dataset page: https://huggingface.co/datasets/SpatiaOS/P3D-Bench.tabulartext-to-3d1K<n<10K2 likes131 downloads3mo agoHugging Face14UKPLab /sparp Dataset Card for Spatial Reasoning Path (SpaRP) Dataset Summary This dataset is a consolidation of SpaRTUN and StepGame datasets with an extension of additional spatial characterization and reasoning path generation. The methodology is explained in our ACL 2024 paper - SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models. The dataset format and fields are normalized across the two… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/sparp.tabular100K<n<1M1 likes130 downloads1y agoHugging Face15asd567557275 /zhtw-roleplay-space-grimoire Space Grimoire RP Corpus (Traditional Chinese) Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0. 中文說明在下方 Dataset Summary Source text 283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.tabulartext-generation10K<n<100K1 likes125 downloads11d agoHugging Face16multimodalart /agent-spaces-tracestabularn<1K0 likes117 downloads5mo agoHugging Face17jamilalthani1 /SPARC SPaRC Dataset Website • Solver • Generator A grid-based puzzle dataset for benchmarking LLMs spatial reasoning capabilities. Data Schema Each record (JSON) includes: id (string): unique puzzle identifier difficulty_level (int) & difficulty_score (float) grid_size: { "height": H, "width": W } polyshapes: JSON string mapping shape IDs to binary grids puzzle_array: 2D array with cell codes (e.g., S, E, +, P-O-112) solution_count (int) and solutions list (with index… See the full description on the dataset page: https://huggingface.co/datasets/jamilalthani1/SPARC.tabular1K<n<10K0 likes113 downloads10mo agoHugging Face18Gradygu3u /spatial-training-full-release-20260604 Spatial Training Full Data Release Full data staging directory for our current Cambrian-P / SSR-style Spatial VLM reproduction work. The directory contains the lightweight reproduction pack plus raw compressed training archives. Local staging uses hardlinks where possible, but upload payload is the full dataset. Size Logical payload: 1076.281 GiB Files: 254 Max single file: 18.0 GiB VSI-590K raw payload: 216.777 GiB Cambrian-S-3M raw payload: 858.004 GiB… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/spatial-training-full-release-20260604.tabularvisual-question-answering1M<n<10M0 likes107 downloads3mo agoHugging Face19jmkey /spatial_mosaic_vqa SpatialMosaic: A Multi-View VLM Dataset for Partial Visibility Description SpatialMosaic is a multi-view visual question answering dataset for evaluating spatial reasoning under partial visibility, occlusion, and low-overlap views. It pairs indoor ScanNet++ and outdoor Waymo scene references with multi-frame VQA annotations. Questions require models to combine fragmented evidence across 2-5 views, rather than answering from a single image. The tasks… See the full description on the dataset page: https://huggingface.co/datasets/jmkey/spatial_mosaic_vqa.tabularvisual-question-answering1K<n<10K0 likes90 downloads3mo agoHugging Face20mmrech /pitvqa-comprehensive-spatial PitVQA Comprehensive Spatial Dataset High-fidelity surgical spatial localization dataset for training vision-language models on pituitary surgery instrument and anatomy detection. 🔗 GitHub: https://github.com/matheus-rech/pit_project 🤖 Trained Model: mmrech/pitvqa-qwen2vl-spatial 📄 Original Dataset: UCL Research Data Repository Dataset Description This dataset contains 10,139 surgical frames with precise spatial annotations for instrument localization and anatomy… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-comprehensive-spatial.tabularvisual-question-answering10K<n<100K1 likes72 downloads8mo agoHugging Face21malaiwah /spark2-5-tiny-fidelity-root-v1 spark2-5 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/spark2-5-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-fidelity-root-v1.tabularn<1K0 likes70 downloads16d agoHugging Face22ri7-5 /spaceship-game-leaderboard Spaceship Game - Leaderboard This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini. Stats Entries: 8 Top Score: 9100 by ri7 Last Updated: 2026-09-21 Published by: ri7-5 Format The leaderboard.json file contains an array of entries: Field Type Description score int Final game score name string Player name date string ISO 8601 timestamp waves_completed int? Number of waves completed Top 10… See the full description on the dataset page: https://huggingface.co/datasets/ri7-5/spaceship-game-leaderboard.tabularothern<1K0 likes67 downloads3d agoHugging Face23DarthBinks /spaceship-game-leaderboard Spaceship Game - Leaderboard This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini. Stats Entries: 2 Top Score: 600 by Pilot Last Updated: 2026-08-25 Published by: DarthBinks Format The leaderboard.json file contains an array of entries: Field Type Description score int Final game score name string Player name date string ISO 8601 timestamp waves_completed int? Number of waves completed… See the full description on the dataset page: https://huggingface.co/datasets/DarthBinks/spaceship-game-leaderboard.tabularothern<1K0 likes65 downloads29d agoHugging Face24ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes56 downloads2mo agoHugging Face25typoverflow /libero_plus_spatial libero_plus_spatial: detailed LeRobot v3.0 This dataset was converted from the LIBERO Plus LeRobot v2.1 libero_plus_spatial partition. The original 8D state and 7D action vectors are preserved exactly as raw_state.ref_state and raw_action.ref_action. Canonical low-dimensional fields follow failure_rollout_data/dataset.md; debug.gripper_eef_* contains the ground-truth next-step relative EEF motion for inspection. Required camera transform for canonical training The… See the full description on the dataset page: https://huggingface.co/datasets/typoverflow/libero_plus_spatial.tabularn<1K0 likes50 downloads2mo agoHugging Face26djdumpling /spatial_reasoningtabularn<1K0 likes48 downloads9mo agoHugging Face27open-llm-leaderboard /arcee-ai__Llama-Spark-detailsgated Dataset Card for Evaluation run of arcee-ai/Llama-Spark Dataset automatically created during the evaluation run of model arcee-ai/Llama-Spark The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Llama-Spark-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face28lidingm /SpatialEvo-160K SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments SpatialEvo-160K Dataset Description SpatialEvo-160K is an offline spatial reasoning QA dataset generated by the Deterministic Geometric Environment (DGE) from SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments. This dataset is not used in the SpatialEvo training pipeline reported in the paper; it is released… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/SpatialEvo-160K.tabularvisual-question-answering100K<n<1M8 likes47 downloads5mo agoHugging Face29spaudwal /BlendNet 📚 BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples. For more details, please visit our GitHub repository or refer to our arXiv paper. 📖 Citation @misc{du2024blenderllmtraininglargelanguage, title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement}, author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/spaudwal/BlendNet.tabular10K<n<100K0 likes47 downloads24d agoHugging Face30aaaded /spark-capacity-boundary-study Capacity-boundary optimization with Spark-X2.5-1.7B This project contains an original evaluation for HER Hack-Astron #6. The report is in DISCUSSION.md. Creating or publishing these artifacts is not an award or payment. The experiment checks whether increasing a six-item 0/1 knapsack's capacity by one causes the model to find the new optimum. Four seeded item families produce eight mathematical instances, each in English and Chinese. Every prompt runs once with thinking off and… See the full description on the dataset page: https://huggingface.co/datasets/aaaded/spark-capacity-boundary-study.tabularn<1K0 likes47 downloads11d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.