datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpIDER-Bench
SpIDER-Bench
Repository dependency graphs for software issue localization — the graph data behind
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
(arXiv:2512.16956).
Each benchmark instance gets one directed multigraph of its repository at the commit the
issue was filed against. Nodes are directories, files, classes and functions carrying
their source; edges are contains / imports / inherits / invokes relations between
them. SpIDER uses these… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SpIDER-Bench.spider_mcqa_v0.2_full
Spider-MCQA
Converted Spider Text-to-SQL (Paper: Yu et al., 2018; HF Dataset) test set into multiple-choice.
The dataset contains 1,034 examples.
Dataset Fields
Each JSON record contains:
query: the schema and natural-language question prompt.
gold_answer: the correct SQL answer.
options: four SQL answer options, including the gold answer and three generated distractors.
correct_option_index: the index of the correct answer in options.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notpaulmartin/spider_mcqa_v0.2_full.spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.doe-genesis-sealed-n2500
Demonstration receipts (n=2500)
These are demonstration envelopes from a filing. n=2500 is a demonstration number. Each row is one trajectory summary, not a 1 kHz pulse and not a timeseries. 15 banks (37,500 rows). Per-row proof_hash. File cryptographic_seal.
What a stranger sees if they cite this zip
They land on a grant-shaped shelf: 15 configs next to each other, mass/μ/booleans/proof_hash. No pulse. No 4×4 taxels. No sentence that this is the direction for… See the full description on the dataset page: https://huggingface.co/datasets/spiderpilot89/doe-genesis-sealed-n2500.spider2-snow-temperature-sweep
Spider 2.0-Snow Temperature-Sweep Rollouts (Qwen3, thinking vs. non-thinking)
Unconstrained language-model rollouts on the Spider 2.0-Snow text-to-SQL benchmark,
sampled from Qwen3 models in both thinking and non-thinking modes. The dataset is
intended for analyzing sampling behavior (temperature, reasoning mode, model size) on hard,
enterprise-scale text-to-SQL. These are generations only — execution correctness (eval)
is added in a separate scoring pass.
Configs… See the full description on the dataset page: https://huggingface.co/datasets/vxef/spider2-snow-temperature-sweep.classical-grasp
Classical grasp (n=2500 envelope + 1 kHz pulse)
Friction cone on a 4×4 pad. Local slip if |τ| > μ N. Micro: outer ring, inner stuck (or shear within 10% of the cone). Macro: inner slip or |v_slip| > 0.005 m/s. Reflex ramps F ← F + scale·dt, clamp 45 N. Law: evaluate_grasp_dynamics in src/physics/dexterous.rs. Reflex twin: ztp_dexterous_evaluate_grasp.
Clock
Envelope rows are summaries of 1 kHz loops. Pulse is 100 steps at dt = 0.001 s.
Envelope rates… See the full description on the dataset page: https://huggingface.co/datasets/spiderpilot89/classical-grasp.MedPriv-Bench_dataset
MedPriv-Bench
MedPriv-Bench evaluates the privacy--utility trade-off of language models in
medical open-ended question answering. Each example contains a synthetic
patient context, injected privacy-sensitive facts, a question, and a
ground-truth answer. The benchmark supports evaluating whether a model remains
clinically useful while avoiding disclosure of the injected facts.
Split sizes
Split
Rows
Composition
train
2,200
1,315 benchmark-construction… See the full description on the dataset page: https://huggingface.co/datasets/Spiderman01/MedPriv-Bench_dataset.SynQL-Spider-Train
Dataset Card for SynQL-Spider-Train
Developed by: Semiotic Labs
Dataset type: [Text-to-SQL]
License: [Apache-2.0]
Dataset Details
Example view of data:
[
{
"question": "What are the names of browsers that have a market share greater than 10% but less than 50%?",
"query": "SELECT name FROM browser WHERE market_share > 10 AND market_share < 50",
"db_id": "browser_web",
"topic_id": "2",
"query_id": "19"
},
...
{… See the full description on the dataset page: https://huggingface.co/datasets/semiotic/SynQL-Spider-Train.fetch_pick_place_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
25
],
"names": [
"s0",
"s1",
"s2",
"s3",
"s4",
"s5",
"s6",
"s7"… See the full description on the dataset page: https://huggingface.co/datasets/SpiderWolf6/fetch_pick_place_v2.fetch_pick_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
25
],
"names": [
"s0",
"s1",
"s2",
"s3",
"s4",
"s5",
"s6",
"s7"… See the full description on the dataset page: https://huggingface.co/datasets/SpiderWolf6/fetch_pick_place.humanoid_pick_stackThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"s0",
"s1",
"s2",
"s3",
"s4",
"s5",
"s6",
"s7"… See the full description on the dataset page: https://huggingface.co/datasets/SpiderWolf6/humanoid_pick_stack.spider-text-to-sqlspider
spider
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
spider-model-outputs-wo-gpt35spider-silk-benchmarkmultilingual-nl2sql-datasets-gen_data_spiderSpider_Redo_Prod
Spider_Redo_Prod
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
spider-model-outputsspider-classifier-training-data
Spider Classifier — Training Manifest
Public release of the training manifest used to fine-tune the
Spiders of New Hampshire species
classifier.
This manifest enumerates every photo used to train, validate, and test the
model. Each row links back to the original observation and photo on
iNaturalist, preserving full attribution and
license metadata.
Source
Model run: 20260528_104624_licensed_dinov2_l_14_reg4_518
Generated: 2026-05-29T02:17:00.691986+00:00… See the full description on the dataset page: https://huggingface.co/datasets/bwirth/spider-classifier-training-data.collabllm-multiturn-spiderfull_filtered_pipeline_base_spiderspider2-snow-embedding-eval
Spider 2.0-SNOW — Grouped Schema-Linking Collection (rich format, v2)
Schema-linking evaluation set built from Spider 2.0-SNOW (547 instances) with
true-shard table/column grouping — collapsing per-year / per-state / per-shard sibling
tables (e.g. EVENTS_20201124 … EVENTS_20210131 → EVENTS_*) and duplicate columns into
single logical nodes, so a 17 883-column mega-schema becomes a tractable retrieval node set.
DETERMINISTIC build (v2). No field requires an LLM call. Column/table… See the full description on the dataset page: https://huggingface.co/datasets/thanhdath/spider2-snow-embedding-eval.
