datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
general-knowledge-mcq-training-pool
General knowledge multiple-choice training pool
Public multiple-choice questions in medicine and health, law, history, philosophy, business and
everyday general knowledge, from four datasets, read at the pinned revisions named below and laid
out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 236665 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/general-knowledge-mcq-training-pool.general_knowledge_booleansft-ready-MuskumPillerum-General-KnowledgeGeneral-Knowledge-VI
📚 Lvoxx/General-Knowledge-VI
Lvoxx/General-Knowledge-VI là bộ dữ liệu kiến thức phổ thông song ngữ (Việt - Anh). Dữ liệu được biên dịch và tối ưu hóa từ bộ dữ liệu gốc MuskumPillerum/General-Knowledge.
Điểm đặc biệt của dataset này là giữ nguyên cặp câu hỏi/trả lời gốc bằng tiếng Anh song song với bản dịch tiếng Việt, phù hợp cho các tác vụ huấn luyện mô hình đa ngôn ngữ hoặc hệ thống RAG đối chiếu.
📋 Mục lục
Cấu trúc dữ liệu
Ví dụ dữ liệu
Cách sử dụng
Nguồn & Ghi… See the full description on the dataset page: https://huggingface.co/datasets/Lvoxx/General-Knowledge-VI.persian-general-knowledge
Dataset Card for persian-gk (Persian General Knowledge)
Dataset Summary
persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training.
Language: Persian (fa)
Size: 5 897 conversations, 2–8 turns… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge.general_knowledge_dataset
Synthetic MMLU CoT
This dataset contains 27,689 synthetic chain-of-thought examples
generated with Qwen/Qwen3-14B on cais/mmlu auxiliary_train
multiple-choice questions.
Columns
question
choices
answer
answer_letter
teacher_output
Provenance and License
The original questions, answer choices, and gold labels come from
cais/mmlu, split auxiliary_train. The
Hugging Face dataset card for cais/mmlu lists its license as
mit. Those source fields retain… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-AttentionSeekers/general_knowledge_dataset.persian-general-knowledge-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset.
The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others.
Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge-cleaned.familicare_health_general_knowledgesept19-general_knowledgegeneral_knowledgegeneral_knowledge_boolean_sample10general_knowledgesept19-general_knowledge-y_trainsept-19-general_knowledge-y_testhealth_general_knowledgegeneral_knowledge_gemini
