datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.GLM-5.2-Finance-80000x
GLM-5.2 · Finance-80000x
80,000x financial related traces distilled from GLM-5.2 on High reasoning
Risk · Markets · Investments · Corporate Finance · Wealth Management
Token Count: 220M
Unique prompts generated with diffusion Gemma-27B answered by GLM-5.2
You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own.
hi - ianncity
GLM-5.2-Logic-Puzzles
GLM-5.2 · Logical Puzzles
6000x traces distilled from GLM-5.2 on High reasoning
Token Count: 5M~?
Distribution:
Puzzles:
•Tokenization blindless ex: counting the r's in strawberry
•Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing)
•Reading comprehension traps
•Temporal reasoning
•Many other categories not worth mentioning
Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.GLM-5.2-Science
GLM-5.2 · Science-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Physics · Chemistry · Biology
Token Count: 160M
Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now
You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own.
hi - ianncity
GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/GLM-5.2-Conversation.synthlabs-GLM-5.2-Science
GLM-5.2 Science Synth Reasoning
Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer.
Dataset Summary
33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source)
33,014 reasoning turns (99.9% format compliance)
Average 3,094 chars per reasoning trace
Models Used
Model
Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.GLM-5.2-Finance-80000x-ru
GLM-5.2-Finance-80000x (RU) — очищенная версия
Набор из 78 652 финансовых reasoning-трассировок на русском языке с цепочками рассуждений <think>. Подготовлен для тонкой настройки (SFT) русскоязычных LLM в финансовой области.
Характеристики
Параметр
Значение
Записей
78 652
Формат
OpenAI chat (messages с role/content)
Структура
строго user → assistant, один ход
Рассуждения <think>…</think>
100% записей
Токенов (GLM-5.2 tokenizer)
~268 млн… See the full description on the dataset page: https://huggingface.co/datasets/mizinovmv/GLM-5.2-Finance-80000x-ru.
