datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.general_knowledge_data
General Knowledge Reproduction Data
This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model.
The corresponding model repository is:
cs-552-2026-catma/general_knowledge_model
The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.deep-general-knowledge-zh
Deep General Knowledge Dialogue Dataset (Chinese)
深度通用知识对话数据集
Dataset Description
High-quality Chinese general knowledge dialogues covering interdisciplinary topics, daily life questions, and miscellaneous knowledge.
高质量中文通用知识对话,涵盖跨学科综合话题、日常生活问题、杂学知识等。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata: Source platform, topic tags… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-general-knowledge-zh.persian-general-knowledge
Dataset Card for persian-gk (Persian General Knowledge)
Dataset Summary
persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training.
Language: Persian (fa)
Size: 5 897 conversations, 2–8 turns… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge.persian-general-knowledge-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset.
The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others.
Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge-cleaned.general_knowledge_dataset
Synthetic MMLU CoT
This dataset contains 27,689 synthetic chain-of-thought examples
generated with Qwen/Qwen3-14B on cais/mmlu auxiliary_train
multiple-choice questions.
Columns
question
choices
answer
answer_letter
teacher_output
Provenance and License
The original questions, answer choices, and gold labels come from
cais/mmlu, split auxiliary_train. The
Hugging Face dataset card for cais/mmlu lists its license as
mit. Those source fields retain… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-AttentionSeekers/general_knowledge_dataset.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/prem7030/General-Knowledge.
