datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
playdate-games
Coding agent session traces for aaaaliou/playdate-games
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/playdate-games.SPADE-Grounding-Corpus-Games-15K
SPADE grounding corpus — games (15k)
Reference documents the SPADE proposer is grounded on when generating cognitive-skill
game environments. 15,000 documents: 10k drawn from a mathematics corpus and 5k from a
science corpus.
Documents
15,000
Setting
games
Fields
Field
Description
text
The document, exactly as embedded in the generation prompt
metadata
domain (mathematics / science) and url (source provenance)
Each generation… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-Games-15K.ru-word-games
Dataset Summary
Dataset contains more than 100k examples of pairs word-description, where description is kind of crossword question. It could be useful for models that generate some description for a word, or try to a guess word from a description.
Source code for parsers and example of project are available here
Key stats:
Number of examples: 133223
Number of sources: 8
Number of unique answers: 35024
subset
count
350_zagadok
350
bashnya_slov
43522
crosswords
39290… See the full description on the dataset page: https://huggingface.co/datasets/artemsnegirev/ru-word-games.steam-games-semanticIds-instructions-v3
Steam Games -- Semantic ID Instruction-Tuning Dataset (v3)
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-8B-Game-semantic-IDs-v3
to reason over the semantic-ID space instead of raw item IDs/embeddings.
Successor to pblrvo/steam-games-semanticIds-instructions
(used for v1/v2), kept as a separate repo rather than overwriting it -- v2's model… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions-v3.steam-games-semanticIds-instructions
Steam Games -- Semantic ID Instruction-Tuning Dataset
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-4B-Game-semantic-IDs to
reason over the semantic-ID space instead of raw item IDs/embeddings.
Train: 299,491 examples (sft_train.jsonl)
Validation: 16,118 examples (sft_val.jsonl)
Special tokens: 1,026 (semantic-ID vocabulary: <|sid_start|>… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions.
