datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpineBench
Dataset Card for SpineBench
Benchmark Details
Paper Information
Benchmark Examples
Benchmark Distribution
Data Format
Data SourceHuman Evaluation of MLLMs Reasoning Performance
Citation
Benchmark Details
SpineBench is a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of LVLM in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical… See the full description on the dataset page: https://huggingface.co/datasets/Silversorrow/SpineBench.spiqa
SPIQA Dataset Card
Dataset Details
Dataset Name: SPIQA (Scientific Paper Image Question Answering)
Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
Github: SPIQA eval and metrics code repo
Dataset Summary: SPIQA is a large-scale and challenging QA dataset focused on figures, tables, and text paragraphs from scientific research papers in various computer science domains. The figures cover a wide variety of plots… See the full description on the dataset page: https://huggingface.co/datasets/google/spiqa.spider-schema
Dataset Card for Spider Schema
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset contains the 166 databases used in the Spider dataset.
Yale Lily Spider Leaderboards
The leaderboard can be seen at https://yale-lily.github.io/spider
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-schema.spider-context-validation
Dataset Card for Spider Context Validation
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to validate spider-fine-tuned LLMs with database context.
Yale Lily Spider Leaderboards
The leaderboard can be seen at https://yale-lily.github.io/spider… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-validation.SpinBench
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
🌐 Project page |
🤗 Dataset |
📑 Paper |
💻 Code
SpinBench is a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision-language models (VLMs).
SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under viewpoint transformation. Since perspective taking requires… See the full description on the dataset page: https://huggingface.co/datasets/YuyouZhang/SpinBench.spider2-litespider-context-instruct
Dataset Card for Spider Context Instruct
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to finetune LLMs in a ### Instruction: and ### Response: format with database context.
Yale Lily Spider Leaderboards
The leaderboard can be seen at… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-context-instruct.spider-realistic
Dataset Card for Spider-Releastic
This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.SPIEval
SPIEval
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
SPIEval is a human-curated benchmark for evaluating whether large language models can act as mobile assistants by proactively retrieving and reasoning over personal information scattered across multiple applications. Given an underspecified user instruction, a model must search structured records, recover the information required for execution, and invoke the appropriate… See the full description on the dataset page: https://huggingface.co/datasets/Junjie-Ye/SPIEval.spivavtor
Dataset Card for Spivavtor
Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model
Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni
Dataset Summary
This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing.
The specific details are as follows:
Task
Examples in Training data
Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.spider-syn
Dataset Card for Sypder-Syn
Spyder-Syn is a human curated variant of the Spider Text-to-SQL database.
The database was created to test the robustness of text-to-SQL models for robustness of synonym substitution.
The source GIT repo for Sypder-Syn is located here: https://github.com/ygan/Spider-Syn
Details regarding the data perterbation methods used and objectives are described in ACL 2021: arXiv
Paper Abstract
Recently, there has been significant progress in… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-syn.spice-circuits-finetune-v5spice-circuits-finetune-v3
SPICE Circuits Fine-Tune V3
A high-quality instruction-following dataset for fine-tuning language models to generate valid, simulation-ready SPICE netlists from natural language descriptions.
Dataset Summary
Property
Value
Total entries
12,471
Format
{"instruction": "...", "output": "..."}
PySpice validation
100% pass
ngspice simulation
99.2% pass (500-entry spot check)
Filepath leaks
0
License
Apache 2.0
What Makes V3… See the full description on the dataset page: https://huggingface.co/datasets/ADI2005/spice-circuits-finetune-v3.spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplesspice-circuits-finetune-v2
SPICE Circuits Fine-tune V2
A clean, validated dataset of 7,410 instruction-output pairs for fine-tuning language models to generate SPICE netlists from natural language descriptions.
Dataset Description
This is Version 2 of the SPICE circuits fine-tuning dataset. V1 was polluted with mixed formats (LTspice, KiCad, standard SPICE) and no validation. V2 is fully validated — every netlist passes PySpice's SpiceParser.build_circuit() gate. No exceptions.… See the full description on the dataset page: https://huggingface.co/datasets/ADI2005/spice-circuits-finetune-v2.spider_12Samples from Spider 1 and Spider 2 for SQLite.
To have DBs locally for Spider 1 refer to the Getting Started of the official website.
All queries here are tested in the databases in test_database.
To have DBs locally for Spider 2 refer to the Quickstart of the github page (the first point is enough)
TaskLAMA
TaskLAMA
This is an unofficial upload of the TaskLAMA data.
TaskLAMA is a novel dataset for Structured Complex Task Decomposition (SCTD).
Some of the data statistics could be found at Spico197/TaskLAMA .
Citation
@misc{yuan2023tasklama,
title={TaskLAMA: Probing the Complex Task Understanding of Language Models},
author={Quan Yuan and Mehran Kazemi and Xin Xu and Isaac Noble and Vaiva Imbrasaite and Deepak Ramachandran},
year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/Spico/TaskLAMA.spider-dpo-1040
Spider DPO 1040
Spider DPO 1040 is a compact Text-to-SQL training dataset for supervised fine-tuning and Direct Preference Optimization. It contains 1,040 preference pairs derived from frontier-model disagreements on Spider V1, plus 7,000 supervised Spider train examples formatted for LLaMA-Factory.
The dataset was created for the companion LoRA adapter jk200201/qwen2.5-coder-7b-sql-dpo.
Important Evaluation Note
The DPO preference pairs in this repository were… See the full description on the dataset page: https://huggingface.co/datasets/jk200201/spider-dpo-1040.spider-natsql-context-instruct
Dataset Card for Spider NatSQL Context Instruct
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset was created to finetune LLMs on the Spider dataset with database context using NatSQL.
NatSQL
NatSQL is an intermediate representation for SQL that simplifies the… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-natsql-context-instruct.spider-text-2-sqlspikeinit-repro-bundle
Reproduction — Training Deep Spiking Neural Networks without Normalization (SpikeInit)
Independent reproduction of ICML 2026 paper #2155 (OpenReview fSV4XWhMB3, Xinyu Shi & Zhaofei Yu)
for the Hugging Face / AlphaXiv ICML-2026 agent-reproduction challenge.
Official code (vendored verbatim at commit b777c7d): https://github.com/xyshi2000/SpikeInit
Claims verified
SpikeInit enables stable training of deep SNNs without BatchNorm. Verified via the init
mechanism:… See the full description on the dataset page: https://huggingface.co/datasets/yaegerknight/spikeinit-repro-bundle.scbe-spine-overlay-proof-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Spine-Overlay Proof v1
Tiny demonstration bundle (18 rows = 3 domains x 6 tongues)
proving that the SCBE 12+ lane code-packet spine handles code,
chemistry, and mechanical motion as overlays on a single tokenizer
substrate, without forking the system.
Why this exists
Every row carries the same baseline lanes (binary, tokenizer, transport,
labels… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-spine-overlay-proof-v1.911-call-transcriptsmascarade-spice-dataset
Ailiance — SPICE & Analog Simulation Q&A
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/mascarade-spice-dataset. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Q&A bilingue (FR/EN) sur la simulation SPICE et l'analyse de circuits analogiques : ngspice, LTspice, modèles MOSFET/BJT, topologies analogiques, ampli-op, filtres actifs, sources de courant, polarisation.
Statistics
Métrique… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-spice-dataset.iso-safety-trajectories
VLA 合成轨迹数据集 — 多参数 · 物理一致 · 免清洗
合成数据天生干净。20 分钟生成 2000 条,真实采集要 44 天。
八大卖点
1. 免清洗 — 省掉 30-40% 项目工时
真实采集数据需要大量人工清洗(去重、去异常、标注校对、格式归一化),通常占项目 30-40% 工时。合成数据天生干净:零噪声、零标注错误、零缺失值。Load 即用。
2. ISO 标准内置 — 合规率 98.7%
ISO 13482 + GB/T 41527 双标准直接嵌入生成逻辑,不是"先生成再检查",而是"从源头保证合规"。合规率 **98.7%**。
3. 三角闭合 — 自洽性 89.75 分
动作-物体-交互三组标签互相校验,不存在"动作对了但物体不对"的不一致数据。自洽性评分 89.75 / 100。
4. 标准化输出 — 100% 格式一致
统一 JSON Schema:51 种动作 × 6… See the full description on the dataset page: https://huggingface.co/datasets/spiralfg/iso-safety-trajectories.sql-create-context-spider-intersectspicy-3.1Airoboros 3.1 dataset with the spicy/decensorship data re-added.
GRAST-SQL-Spider
⚠️ This dataset has been deprecated. Please use the updated version below, which includes improved quality checks:https://huggingface.co/datasets/griffith-bigdata/GRAST-SL-evaluation-set
GRAST-SQL Spider Training & Evaluation Dataset
This dataset is processed from the original Spider dataset, with extracted schema information and used_columns from SQL queries. It is prepared solely for training, and evaluating schema filtering in the context of the GRAST-SQL paper.… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/GRAST-SQL-Spider.spider-queries-trainSPIQA_50K_Re
SPIQA 50K Re-annotated
QA annotations on scientific paper figures from the SPIQA dataset.
Dataset Structure
Each sample contains:
Field
Description
image
Relative path to the figure image
question
Question about the figure
thinking
Chain-of-thought reasoning
answer
Final answer
Splits
train (spiqa_50k.json): 50,000 samples
train_reannotate (spiqa_50k_reannotate.json): 49,975 samples with more detailed chain-of-thought reasoning… See the full description on the dataset page: https://huggingface.co/datasets/scimdr/SPIQA_50K_Re.
