datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.generalization-dynamics-evals
Generalization Dynamics — Main Eval Suite
Prepared test sets for the 6 main evaluation families from
Generalization dynamics across fine-tuning
(Table 1).
Use with the unified runner:
https://github.com/jiaxin-wen/FT-generalization/tree/main/release
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")
Or browse a single task (the dataset viewer shows all configs):
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.General365_Public
🧩 General365: Benchmarking General Reasoning in LLMs Across Diverse and Challenging Tasks
📃 Paper • 🌐 Project Page • 🏆 Leaderboard •
💻 Github
📖 Introduction
We present General365, a highly challenging and diverse benchmark for evaluating the general reasoning capabilities in LLMs.
"General Reasoning" refers to reasoning tasks that depend exclusively on general knowledge.
We define general knowledge as knowledge within the K-12 scope (such as common sense… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/General365_Public.dataset-ohada-droit-commercial-general-echantillon
Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.gujarati-general-purpose-instruction
Gujarati General-Purpose Instruction Dataset (GGJI v1)
Dataset Summary
GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.general-knowledge-mcq-training-pool
General knowledge multiple-choice training pool
Public multiple-choice questions in medicine and health, law, history, philosophy, business and
everyday general knowledge, from four datasets, read at the pinned revisions named below and laid
out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 236665 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/general-knowledge-mcq-training-pool.General-Evol-VQA
Dataset Card for General-Evol-VQA-1.2M
This dataset has been carefully curated to enhance the general instruction capabilities of Vision-Language Models (VLMs). It comprises two subsets:
600k English samples
600k Korean samples
We recommend using this dataset alongside other task-specific datasets (e.g., OCR, Language, code, math, ...) to improve performance and achieve more robust model capabilities.
Made by: maum.ai Brain NLP. Jaeyoon Jung, Yoonshik Kim
Dataset Target… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/General-Evol-VQA.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.General365_Public
🧩 General365: Benchmarking General Reasoning in LLMs Across Diverse and Challenging Tasks
📃 Paper • 🌐 Project Page • 🏆 Leaderboard •
💻 Github
📖 Introduction
We present General365, a highly challenging and diverse benchmark for evaluating the general reasoning capabilities in LLMs.
"General Reasoning" refers to reasoning tasks that depend exclusively on general knowledge.
We define general knowledge as knowledge within the K-12 scope (such as common sense… See the full description on the dataset page: https://huggingface.co/datasets/Ckriman/General365_Public.GeneralTextCorpus
Mixed Content Dataset
Description:This dataset contains a diverse collection of text from multiple domains, including general knowledge, cooking, articles, and more. Each entry typically includes text content along with metadata such as source, title, and language.
The dataset is structured to support research, analysis, or training of NLP models on varied textual content.
Data Structure:Each item typically contains:
id: Unique identifier
text: Main text content
meta: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/ademchaoua/GeneralTextCorpus.vi_instruct_general_dataset_cleaned
Vietnamese Instruct General Dataset (Cleaned & ShareGPT format)
Dataset Description
This dataset is a cleaned version of VTSNLP/instruct_general_dataset. It has been specifically mapped to the ShareGPT format to be readily compatible with fine-tuning frameworks such as Unsloth, Axolotl, and LLaMA-Factory.
Format
The dataset uses the standard ShareGPT structure. Each row contains a conversations list with human and gpt turns, alongside a meta… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vi_instruct_general_dataset_cleaned.general_knowledge_dataset
General Knowledge SFT Dataset
This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning.
The dataset has two splits.
Split
Rows
Purpose
train
26,120
LoRA SFT training split
valid
2,000
LoRA SFT validation split
Sources
The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.GeneralHistoryOfAfricaXI
GeneralHistoryOfAfricaXI
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
3731
Avg chars/chunk
696
Avg images/chunk
0.00
Source files
2
Duplicates removed
1
Quality filtered
70
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned text… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/GeneralHistoryOfAfricaXI.qulture-general-knowledge-dataset
Qulture General Knowledge Question Dataset
An open dataset containing 16 families of general knowledge questions, for a total of 64 records.
General-Knowledge-VI
📚 Lvoxx/General-Knowledge-VI
Lvoxx/General-Knowledge-VI là bộ dữ liệu kiến thức phổ thông song ngữ (Việt - Anh). Dữ liệu được biên dịch và tối ưu hóa từ bộ dữ liệu gốc MuskumPillerum/General-Knowledge.
Điểm đặc biệt của dataset này là giữ nguyên cặp câu hỏi/trả lời gốc bằng tiếng Anh song song với bản dịch tiếng Việt, phù hợp cho các tác vụ huấn luyện mô hình đa ngôn ngữ hoặc hệ thống RAG đối chiếu.
📋 Mục lục
Cấu trúc dữ liệu
Ví dụ dữ liệu
Cách sử dụng
Nguồn & Ghi… See the full description on the dataset page: https://huggingface.co/datasets/Lvoxx/General-Knowledge-VI.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/histoire-general-afrique-global-adaption.Inspector-General-Act-of-1978
Dataset Description
The Inspector General Act of 1978 Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the original statutory framework used to establish independent Offices of Inspector General within selected executive-branch departments and agencies.
The dataset was developed from the enacted text of the Inspector General Act of 1978, Public Law 95-452, 92 Stat. 1101, approved on October 12, 1978. The statute… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Inspector-General-Act-of-1978.hindi_eval_general_mcqoffences_and_penalties_in_general_2018_datasetgeneral_knowledge_benchmark
General Knowledge Benchmark Splits
This dataset contains the held-out benchmark splits used for offline model selection and evaluation of the MNLP general knowledge specialist.
These benchmarks were not used for LoRA SFT training. The SFT train and validation splits are stored separately in:
cs-552-2026-databand/general_knowledge_dataset
Splits
Split
Rows
Sampling strategy
Coverage
mmlu_pro
2,000
Uniform across categories
Robust multi-task knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_benchmark.General_Conversation_Mixed_Datasetpersian-general-knowledge
Dataset Card for persian-gk (Persian General Knowledge)
Dataset Summary
persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training.
Language: Persian (fa)
Size: 5 897 conversations, 2–8 turns… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge.General_Conversation_Mixed_Datasetsota-generalbash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
