datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPADE-Environments-Qwen3-30B-Games
SPADE generated environments: games
Paper | Code | All artifacts
Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each.
Environments
3310
Training steps covered
113 (step 0 to 396)
With skill label
3119
Designer / agent model
Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.danbooru-tag-csv
danbooru-tag-csv
CSV files of Danbooru tags.
Dataset Description
This project manages CSVs of Danbooru tags, which can be used by Danbooru related applications and libraries.
Dataset Creation
These CSV files were created using the following datasets:
itterative/danbooru_wikis_full
trojblue/danbooru2025-metadata
License
This dataset is released under the MIT License.
SPADE-Environment-Pool-GPT5.5-Games
SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games
This public dataset contains 7,872 validated Python game environments for actor-only SPARE training.
Six cognitive skills, exactly 1,312 environments per skill
Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl
Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0
Maximum 25 turns and 32K generation context
Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.SPADES
SPADES dataset
SPADES - SPAcecraft Pose Estimation Dataset using Event Sensing, a unique and new space dataset designed to advance spacecraft pose estimation research. SPADES dataset contains two categories of data: Synthetic and Real.
Synthetic dataset focuses on simulating RGB images of a satellite target—in this case, Proba-2—by moving a spacecraft model along predefined trajectories within the simulator’s camera field of view. To generate realistic imagery, the Unreal… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPADES.SPADE-Environments-ToolUse
SPADE generated environments: tool use
Paper | Code | All artifacts
Multi-turn tool-use environments written by the SPADE designer during training, pooled
across every captured run. 2,231 environments across 7 runs and two model scales (30B-A3B and 4B).
Source run
Scale
Environments
qwen3-30b-0617-tooluse-regen32-mixed
30B-A3B
41
qwen3-30b-0624-tooluse-blend
30B-A3B
243
qwen3-30b-0703-tooluse-glory-kl005
30B-A3B
260
qwen3-4b-0630-tooluse-eval-aligned-r32
4B
456… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse.SPADES-RGBSPADE-Environment-Pool-GPT5.5-ToolUse
SPARE GPT-5.5 Multi-Turn Tool-Use Games v1
A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training.
Training alignment
Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration
400 rollouts x 24 games/rollout = 9,600 no-reuse games required
11,039 validated games provide 1,439 games of headroom
Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.GovReport_Corpus_1024_2040
GovReport MIA fine-tuning corpus
This card documents the split layout for
spadeMIA/GovReport_Corpus_1024_2040,
a GovReport corpus used for membership-inference attack (MIA) experiments on
fine-tuned language models.
Membership labels are defined relative to the fine-tuning population.
train is the only split used for fine-tuning, and every row has label = 1.
test remains held out, and every row has label = 0. evaluation is a fixed,
balanced MIA candidate set. Fine-tune on train… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/GovReport_Corpus_1024_2040.GoodWiki_Corpus_1024_2040
GoodWiki 1024–2040: paragraph-truncated MIA fine-tuning corpus
A deterministic, paragraph-truncated corpus of English Wikipedia Good/Featured
articles, built from euirim/goodwiki
for membership inference attack (MIA) experiments on fine-tuned language models.
Membership labels are defined relative to the fine-tuning population.
train (10,000 rows) is the only split used for fine-tuning, and every row has
label = 1. test (1,000 rows) remains held out, and every row has label =… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/GoodWiki_Corpus_1024_2040.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.SPADE-Grounding-Corpus-ToolUse-15K
SPADE grounding corpus: tool use (15k)
Reference documents the SPADE Environment Designer is grounded on when generating multi-turn tool-use environments. 15,552 source files drawn from nvidia/Nemotron-Pretraining-Code-v3.
Documents
15,552
Setting
tool_use
Fields
text (the document), metadata (source provenance)
Each generation prompt embeds one sampled document, so the environments a Designer
writes stay anchored to a real concept or technique rather than… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-ToolUse-15K.SPADE-Grounding-Corpus-Games-15K
SPADE grounding corpus — games (15k)
Reference documents the SPADE proposer is grounded on when generating cognitive-skill
game environments. 15,000 documents: 10k drawn from a mathematics corpus and 5k from a
science corpus.
Documents
15,000
Setting
games
Fields
Field
Description
text
The document, exactly as embedded in the generation prompt
metadata
domain (mathematics / science) and url (source provenance)
Each generation… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-Games-15K.minimind-o_dataset
For full documentation, please refer to the GitHub repository: https://github.com/jingyaogong/minimind-o
Technical Report: https://arxiv.org/abs/2605.03937
SPADE-Environments-Qwen3-30B-ToolUse
qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005 — generated environments
Environments generated by the SPARE proposer during training run
2hjdrbeh (qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
260
Steps covered
7 (step 0–192)
With recovered skill
260
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/SPADE-Environments-Qwen3-30B-ToolUse.SpaDE_datagovreport_finetune_corpuspmc_finetune_corpus_1024-2040_tokens
PMC 1024-2040 Biomedical Fine-Tuning Corpus
Summary
This is a cleaned biomedical long-text corpus for autoregressive language-model
fine-tuning and held-out evaluation.
split
rows
role
train
10,000
fine-tuning
test
1,000
held-out evaluation
The public schema is text-only:
text: string
No PMCID, date, license, URL, or provenance fields are included in the public
dataset files.
Token Contract
The corpus is built for… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/pmc_finetune_corpus_1024-2040_tokens.arxiv_finetune_corpusgovreport_evaluation_benchmarkarxiv_benchmark_longspade
SPADE: A Multilingual Dataset for Speech Partial Deepfake Detection and Localization
The repo contains the SPADE dataset, containing partially edited speech deepfakes in 12 languages, generated by up to five systems per language.
Subsets are either named {language}-{model}-edited or {language}-original.
Edited and original audio are stored in separate subsets, so they can be downloaded independently.
For each edited audio file, the original file corresponding to it can be found… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/spade.pubmed_finetune_corpusspade-100proteins
SPADE 100-protein dataset
Dataset used to evaluate SPADE in the paper SPADE: Fast Drug Discovery by
Learning from Sparse Data. It is a subset of a larger 1.5M-entry
PubChem-derived ligand-protein affinity dataset (full release to follow),
restricted to the 100 proteins on which SPADE and all baselines are
benchmarked in the main results (Sec. 4 of the paper).
Files
File
Shape
Description
proteins.csv
100 rows
index_protein, uniprot_id, sequence (FASTA)… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-99/spade-100proteins.govreport_benchmark_longpubmed_benchmark_longgovreport_finetune_corpus_1024-2040_tokens_2000samplepubmed_preprocessedSPADESspade-sae-fullSpade
