CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Silversorrow /SpineBench Dataset Card for SpineBench Benchmark Details Paper Information Benchmark Examples Benchmark Distribution Data Format Data SourceHuman Evaluation of MLLMs Reasoning Performance Citation Benchmark Details SpineBench is a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of LVLM in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical… See the full description on the dataset page: https://huggingface.co/datasets/Silversorrow/SpineBench.imagevisual-question-answering10K<n<100K0 likes10k downloads11mo agoHugging Face02google /spiqa SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific&nbsp;Paper&nbsp;Image&nbsp;Question&nbsp;Answering) Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers Github: SPIQA eval and metrics code repo Dataset Summary: SPIQA is a large-scale and challenging QA dataset focused on figures, tables, and text paragraphs from scientific research papers in various computer science domains. The figures cover a wide variety of plots… See the full description on the dataset page: https://huggingface.co/datasets/google/spiqa.textquestion-answeringn<1K48 likes1.3k downloads2y agoHugging Face03richardr1126 /spider-schema Dataset Card for Spider Schema Dataset Summary Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases. This dataset contains the 166 databases used in the Spider dataset. Yale Lily Spider Leaderboards The leaderboard can be seen at https://yale-lily.github.io/spider Languages The text in… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-schema.textn<1K8 likes639 downloads3y agoHugging Face04richardr1126 /spider-context-validation Dataset Card for Spider Context Validation Dataset Summary Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases. This dataset was created to validate spider-fine-tuned LLMs with database context. Yale Lily Spider Leaderboards The leaderboard can be seen at https://yale-lily.github.io/spider… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-validation.text1K<n<10K0 likes365 downloads3y agoHugging Face05YuyouZhang /SpinBench SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs 🌐 Project page | 🤗 Dataset | 📑 Paper | 💻 Code SpinBench is a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision-language models (VLMs). SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under viewpoint transformation. Since perspective taking requires… See the full description on the dataset page: https://huggingface.co/datasets/YuyouZhang/SpinBench.imagequestion-answering1K<n<10K5 likes249 downloads6mo agoHugging Face06xlangai /spider2-litetextn<1K12 likes213 downloads2y agoHugging Face07richardr1126 /spider-context-instruct Dataset Card for Spider Context Instruct Dataset Summary Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases. This dataset was created to finetune LLMs in a ### Instruction: and ### Response: format with database context. Yale Lily Spider Leaderboards The leaderboard can be seen at… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-instruct.text1K<n<10K1 likes157 downloads3y agoHugging Face08aherntech /spider-realistic Dataset Card for Spider-Releastic This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.textn<1K3 likes157 downloads3y agoHugging Face09Junjie-Ye /SPIEval SPIEval SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPIEval is a human-curated benchmark for evaluating whether large language models can act as mobile assistants by proactively retrieving and reasoning over personal information scattered across multiple applications. Given an underspecified user instruction, a model must search structured records, recover the information required for execution, and invoke the appropriate… See the full description on the dataset page: https://huggingface.co/datasets/Junjie-Ye/SPIEval.texttext-generationn<1K1 likes154 downloads2mo agoHugging Face10grammarly /spivavtor Dataset Card for Spivavtor Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni Dataset Summary This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing. The specific details are as follows: Task Examples in Training data Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.texttext-generation10K<n<100K5 likes125 downloads2y agoHugging Face11aherntech /spider-syn Dataset Card for Sypder-Syn Spyder-Syn is a human curated variant of the Spider Text-to-SQL database. The database was created to test the robustness of text-to-SQL models for robustness of synonym substitution. The source GIT repo for Sypder-Syn is located here: https://github.com/ygan/Spider-Syn Details regarding the data perterbation methods used and objectives are described in ACL 2021: arXiv Paper Abstract Recently, there has been significant progress in… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-syn.text1K<n<10K1 likes121 downloads3y agoHugging Face12ADI2005 /spice-circuits-finetune-v5text10K<n<100K0 likes94 downloads28d agoHugging Face13ADI2005 /spice-circuits-finetune-v3 SPICE Circuits Fine-Tune V3 A high-quality instruction-following dataset for fine-tuning language models to generate valid, simulation-ready SPICE netlists from natural language descriptions. Dataset Summary Property Value Total entries 12,471 Format {"instruction": "...", "output": "..."} PySpice validation 100% pass ngspice simulation 99.2% pass (500-entry spot check) Filepath leaks 0 License Apache 2.0 What Makes V3… See the full description on the dataset page: https://huggingface.co/datasets/ADI2005/spice-circuits-finetune-v3.text10K<n<100K0 likes77 downloads1mo agoHugging Face14RiddleHe /spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplestextn<1K0 likes70 downloads10mo agoHugging Face15ADI2005 /spice-circuits-finetune-v2 SPICE Circuits Fine-tune V2 A clean, validated dataset of 7,410 instruction-output pairs for fine-tuning language models to generate SPICE netlists from natural language descriptions. Dataset Description This is Version 2 of the SPICE circuits fine-tuning dataset. V1 was polluted with mixed formats (LTspice, KiCad, standard SPICE) and no validation. V2 is fully validated — every netlist passes PySpice's SpiceParser.build_circuit() gate. No exceptions.… See the full description on the dataset page: https://huggingface.co/datasets/ADI2005/spice-circuits-finetune-v2.texttext-generation1K<n<10K0 likes65 downloads1mo agoHugging Face16sarus-tech /spider_12Samples from Spider 1 and Spider 2 for SQLite. To have DBs locally for Spider 1 refer to the Getting Started of the official website. All queries here are tested in the databases in test_database. To have DBs locally for Spider 2 refer to the Quickstart of the github page (the first point is enough) text10K<n<100K0 likes60 downloads2y agoHugging Face17Spico /TaskLAMA TaskLAMA This is an unofficial upload of the TaskLAMA data. TaskLAMA is a novel dataset for Structured Complex Task Decomposition (SCTD). Some of the data statistics could be found at Spico197/TaskLAMA . Citation @misc{yuan2023tasklama, title={TaskLAMA: Probing the Complex Task Understanding of Language Models}, author={Quan Yuan and Mehran Kazemi and Xin Xu and Isaac Noble and Vaiva Imbrasaite and Deepak Ramachandran}, year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/Spico/TaskLAMA.text1K<n<10K2 likes58 downloads3y agoHugging Face18jk200201 /spider-dpo-1040 Spider DPO 1040 Spider DPO 1040 is a compact Text-to-SQL training dataset for supervised fine-tuning and Direct Preference Optimization. It contains 1,040 preference pairs derived from frontier-model disagreements on Spider V1, plus 7,000 supervised Spider train examples formatted for LLaMA-Factory. The dataset was created for the companion LoRA adapter jk200201/qwen2.5-coder-7b-sql-dpo. Important Evaluation Note The DPO preference pairs in this repository were… See the full description on the dataset page: https://huggingface.co/datasets/jk200201/spider-dpo-1040.texttext-generation1K<n<10K2 likes58 downloads3mo agoHugging Face19richardr1126 /spider-natsql-context-instruct Dataset Card for Spider NatSQL Context Instruct Dataset Summary Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases. This dataset was created to finetune LLMs on the Spider dataset with database context using NatSQL. NatSQL NatSQL is an intermediate representation for SQL that simplifies the… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-natsql-context-instruct.text1K<n<10K0 likes56 downloads3y agoHugging Face20hujudev /spider-text-2-sqltext1K<n<10K0 likes55 downloads2y agoHugging Face21yaegerknight /spikeinit-repro-bundle Reproduction — Training Deep Spiking Neural Networks without Normalization (SpikeInit) Independent reproduction of ICML 2026 paper #2155 (OpenReview fSV4XWhMB3, Xinyu Shi & Zhaofei Yu) for the Hugging Face / AlphaXiv ICML-2026 agent-reproduction challenge. Official code (vendored verbatim at commit b777c7d): https://github.com/xyshi2000/SpikeInit Claims verified SpikeInit enables stable training of deep SNNs without BatchNorm. Verified via the init mechanism:… See the full description on the dataset page: https://huggingface.co/datasets/yaegerknight/spikeinit-repro-bundle.textn<1K1 likes54 downloads2mo agoHugging Face22issdandavis /scbe-spine-overlay-proof-v1 Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data. SCBE Spine-Overlay Proof v1 Tiny demonstration bundle (18 rows = 3 domains x 6 tongues) proving that the SCBE 12+ lane code-packet spine handles code, chemistry, and mechanical motion as overlays on a single tokenizer substrate, without forking the system. Why this exists Every row carries the same baseline lanes (binary, tokenizer, transport, labels… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-spine-overlay-proof-v1.textn<1K0 likes53 downloads2mo agoHugging Face23spikecodes /911-call-transcriptstextn<1K3 likes52 downloads2y agoHugging Face24Ailiance-fr /mascarade-spice-dataset Ailiance — SPICE & Analog Simulation Q&A 🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/mascarade-spice-dataset. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025). Q&A bilingue (FR/EN) sur la simulation SPICE et l'analyse de circuits analogiques : ngspice, LTspice, modèles MOSFET/BJT, topologies analogiques, ampli-op, filtres actifs, sources de courant, polarisation. Statistics Métrique… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-spice-dataset.texttext-generation1K<n<10K0 likes52 downloads5mo agoHugging Face25spiralfg /iso-safety-trajectories VLA 合成轨迹数据集 — 多参数 · 物理一致 · 免清洗 合成数据天生干净。20 分钟生成 2000 条,真实采集要 44 天。 八大卖点 1. 免清洗 — 省掉 30-40% 项目工时 真实采集数据需要大量人工清洗(去重、去异常、标注校对、格式归一化),通常占项目 30-40% 工时。合成数据天生干净:零噪声、零标注错误、零缺失值。Load 即用。 2. ISO 标准内置 — 合规率 98.7% ISO 13482 + GB/T 41527 双标准直接嵌入生成逻辑,不是"先生成再检查",而是"从源头保证合规"。合规率 **98.7%**。 3. 三角闭合 — 自洽性 89.75 分 动作-物体-交互三组标签互相校验,不存在"动作对了但物体不对"的不一致数据。自洽性评分 89.75 / 100。 4. 标准化输出 — 100% 格式一致 统一 JSON Schema:51 种动作 × 6… See the full description on the dataset page: https://huggingface.co/datasets/spiralfg/iso-safety-trajectories.textroboticsn<1K0 likes52 downloads2mo agoHugging Face26samlhuillier /sql-create-context-spider-intersecttext1K<n<10K0 likes50 downloads3y agoHugging Face27unalignment /spicy-3.1Airoboros 3.1 dataset with the spicy/decensorship data re-added. text100K<n<1M29 likes45 downloads3y agoHugging Face28griffith-bigdata /GRAST-SQL-Spider ⚠️ This dataset has been deprecated. Please use the updated version below, which includes improved quality checks:https://huggingface.co/datasets/griffith-bigdata/GRAST-SL-evaluation-set GRAST-SQL Spider Training & Evaluation Dataset This dataset is processed from the original Spider dataset, with extracted schema information and used_columns from SQL queries. It is prepared solely for training, and evaluating schema filtering in the context of the GRAST-SQL paper.… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/GRAST-SQL-Spider.text1K<n<10K0 likes45 downloads1mo agoHugging Face29target-benchmark /spider-queries-traintext1K<n<10K0 likes41 downloads2y agoHugging Face30scimdr /SPIQA_50K_Re SPIQA 50K Re-annotated QA annotations on scientific paper figures from the SPIQA dataset. Dataset Structure Each sample contains: Field Description image Relative path to the figure image question Question about the figure thinking Chain-of-thought reasoning answer Final answer Splits train (spiqa_50k.json): 50,000 samples train_reannotate (spiqa_50k_reannotate.json): 49,975 samples with more detailed chain-of-thought reasoning… See the full description on the dataset page: https://huggingface.co/datasets/scimdr/SPIQA_50K_Re.textquestion-answering10K<n<100K1 likes41 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.