datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
steam-games-semanticIds-instructions-v3
Steam Games -- Semantic ID Instruction-Tuning Dataset (v3)
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-8B-Game-semantic-IDs-v3
to reason over the semantic-ID space instead of raw item IDs/embeddings.
Successor to pblrvo/steam-games-semanticIds-instructions
(used for v1/v2), kept as a separate repo rather than overwriting it -- v2's model… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions-v3.steam-reviews-scraper
Steam Reviews Scraper · Game Reviews, Ratings & Playtime
Scrape Steam game reviews, ratings, playtime, helpfulness votes, and purchase types across any Steam App ID. Fast HTTP scraper, no login required.
Rows in this dataset
9,260
Fields
26
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a keyword list: a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/steam-reviews-scraper.steam-games-semanticIds-instructions
Steam Games -- Semantic ID Instruction-Tuning Dataset
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-4B-Game-semantic-IDs to
reason over the semantic-ID space instead of raw item IDs/embeddings.
Train: 299,491 examples (sft_train.jsonl)
Validation: 16,118 examples (sft_val.jsonl)
Special tokens: 1,026 (semantic-ID vocabulary: <|sid_start|>… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions.steam_rec_system
Steam 游戏推荐系统数据集
本数据集包含用于训练游戏推荐系统的对话格式数据,主要涵盖以下几个方面:
游戏属性知识问答:
多选题形式 (qa_choice_8)
真伪判断任务 (qa_judge)
用户行为建模:
序列回忆任务 (seqrec_next_item)
排序任务 (seqrec_rank)
偏好预测任务 (seqrec_preference)
数据集采用对话格式,适用于大语言模型的领域微调。
数据格式
{
"conversations": [
{
"from": "human",
"value": "问题内容"
},
{
"from": "gpt",
"value": "回答内容"
}
],
"system": "You are a Steam game expert assistant.",
"field": "任务类型"
}
统计信息
总数据量:156,598条
cards-steammachine
