CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes47k downloads7mo agoHugging Face02Scale-or-Reason /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.textquestion-answering1M<n<10M6 likes1.2k downloads3mo agoHugging Face03Sidsidney /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.textquestion-answering1M<n<10M4 likes764 downloads10mo agoHugging Face04MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes549 downloads10mo agoHugging Face05jiaxin-wen /generalization-dynamics-evals Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.texttext-classification10K<n<100K0 likes181 downloads4mo agoHugging Face06HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes144 downloads10mo agoHugging Face07meituan-longcat /General365_Public 🧩 General365: Benchmarking General Reasoning in LLMs Across Diverse and Challenging Tasks 📃 Paper • 🌐 Project Page • 🏆 Leaderboard • 💻 Github 📖 Introduction We present General365, a highly challenging and diverse benchmark for evaluating the general reasoning capabilities in LLMs. "General Reasoning" refers to reasoning tasks that depend exclusively on general knowledge. We define general knowledge as knowledge within the K-12 scope (such as common sense… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/General365_Public.textquestion-answeringn<1K10 likes137 downloads6mo agoHugging Face08Bisilivan /dataset-ohada-droit-commercial-general-echantillon Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon Description Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.texttext-generationn<1K1 likes86 downloads2mo agoHugging Face09tkdonda /gujarati-general-purpose-instruction Gujarati General-Purpose Instruction Dataset (GGJI v1) Dataset Summary GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.texttext-generation10K<n<100K0 likes79 downloads2mo agoHugging Face10Emulated-Inc /general-knowledge-mcq-training-pool General knowledge multiple-choice training pool Public multiple-choice questions in medicine and health, law, history, philosophy, business and everyday general knowledge, from four datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 236665 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/general-knowledge-mcq-training-pool.textquestion-answering100K<n<1M0 likes66 downloads13d agoHugging Face11maum-ai /General-Evol-VQA Dataset Card for General-Evol-VQA-1.2M This dataset has been carefully curated to enhance the general instruction capabilities of Vision-Language Models (VLMs). It comprises two subsets: 600k English samples 600k Korean samples We recommend using this dataset alongside other task-specific datasets (e.g., OCR, Language, code, math, ...) to improve performance and achieve more robust model capabilities. Made by: maum.ai Brain NLP. Jaeyoon Jung, Yoonshik Kim Dataset Target… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/General-Evol-VQA.imagevisual-question-answering1M<n<10M5 likes58 downloads2y agoHugging Face12MaatAI /histoire-general-afrique-global-adaption This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Svngoku/Histoire-General-Afrique-Global This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.textquestion-answering1K<n<10K1 likes54 downloads5mo agoHugging Face13Ckriman /General365_Public 🧩 General365: Benchmarking General Reasoning in LLMs Across Diverse and Challenging Tasks 📃 Paper • 🌐 Project Page • 🏆 Leaderboard • 💻 Github 📖 Introduction We present General365, a highly challenging and diverse benchmark for evaluating the general reasoning capabilities in LLMs. "General Reasoning" refers to reasoning tasks that depend exclusively on general knowledge. We define general knowledge as knowledge within the K-12 scope (such as common sense… See the full description on the dataset page: https://huggingface.co/datasets/Ckriman/General365_Public.textquestion-answeringn<1K0 likes52 downloads5mo agoHugging Face14ademchaoua /GeneralTextCorpus Mixed Content Dataset Description:This dataset contains a diverse collection of text from multiple domains, including general knowledge, cooking, articles, and more. Each entry typically includes text content along with metadata such as source, title, and language. The dataset is structured to support research, analysis, or training of NLP models on varied textual content. Data Structure:Each item typically contains: id: Unique identifier text: Main text content meta: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/ademchaoua/GeneralTextCorpus.texttext-generation10K<n<100K0 likes50 downloads9mo agoHugging Face15vlinhd11 /vi_instruct_general_dataset_cleaned Vietnamese Instruct General Dataset (Cleaned & ShareGPT format) Dataset Description This dataset is a cleaned version of VTSNLP/instruct_general_dataset. It has been specifically mapped to the ShareGPT format to be readily compatible with fine-tuning frameworks such as Unsloth, Axolotl, and LLaMA-Factory. Format The dataset uses the standard ShareGPT structure. Each row contains a conversations list with human and gpt turns, alongside a meta… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vi_instruct_general_dataset_cleaned.textquestion-answering1M<n<10M0 likes49 downloads21d agoHugging Face16cs-552-2026-databand /general_knowledge_dataset General Knowledge SFT Dataset This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning. The dataset has two splits. Split Rows Purpose train 26,120 LoRA SFT training split valid 2,000 LoRA SFT validation split Sources The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.textquestion-answering10K<n<100K0 likes46 downloads4mo agoHugging Face17miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes43 downloads1y agoHugging Face18Svngoku /GeneralHistoryOfAfricaXI GeneralHistoryOfAfricaXI Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline. Dataset Summary Metric Value Total chunks 3731 Avg chars/chunk 696 Avg images/chunk 0.00 Source files 2 Duplicates removed 1 Quality filtered 70 Schema Column Type Description chunk_id string Unique identifier: filename_chunk_N text string Raw markdown chunk with image refs text_clean string Cleaned text… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/GeneralHistoryOfAfricaXI.imagetext-generation1K<n<10K1 likes43 downloads3mo agoHugging Face19Qulture /qulture-general-knowledge-dataset Qulture General Knowledge Question Dataset An open dataset containing 16 families of general knowledge questions, for a total of 64 records. tabularquestion-answeringn<1K0 likes34 downloads1mo agoHugging Face20Lvoxx /General-Knowledge-VI 📚 Lvoxx/General-Knowledge-VI Lvoxx/General-Knowledge-VI là bộ dữ liệu kiến thức phổ thông song ngữ (Việt - Anh). Dữ liệu được biên dịch và tối ưu hóa từ bộ dữ liệu gốc MuskumPillerum/General-Knowledge. Điểm đặc biệt của dataset này là giữ nguyên cặp câu hỏi/trả lời gốc bằng tiếng Anh song song với bản dịch tiếng Việt, phù hợp cho các tác vụ huấn luyện mô hình đa ngôn ngữ hoặc hệ thống RAG đối chiếu. 📋 Mục lục Cấu trúc dữ liệu Ví dụ dữ liệu Cách sử dụng Nguồn & Ghi… See the full description on the dataset page: https://huggingface.co/datasets/Lvoxx/General-Knowledge-VI.textquestion-answering10K<n<100K1 likes32 downloads8mo agoHugging Face21Svngoku /histoire-general-afrique-global-adaption This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Svngoku/Histoire-General-Afrique-Global This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/histoire-general-afrique-global-adaption.textquestion-answering1K<n<10K1 likes28 downloads5mo agoHugging Face22leeroy-jankins /Inspector-General-Act-of-1978 Dataset Description The Inspector General Act of 1978 Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the original statutory framework used to establish independent Offices of Inspector General within selected executive-branch departments and agencies. The dataset was developed from the enacted text of the Inspector General Act of 1978, Public Law 95-452, 92 Stat. 1101, approved on October 12, 1978. The statute… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Inspector-General-Act-of-1978.documentquestion-answeringn<1K1 likes28 downloads3mo agoHugging Face23cmeraki /hindi_eval_general_mcqtextquestion-answering1K<n<10K3 likes25 downloads3y agoHugging Face24iradukunda-dev /offences_and_penalties_in_general_2018_datasettexttext-classificationn<1K0 likes25 downloads9mo agoHugging Face25cs-552-2026-databand /general_knowledge_benchmark General Knowledge Benchmark Splits This dataset contains the held-out benchmark splits used for offline model selection and evaluation of the MNLP general knowledge specialist. These benchmarks were not used for LoRA SFT training. The SFT train and validation splits are stored separately in: cs-552-2026-databand/general_knowledge_dataset Splits Split Rows Sampling strategy Coverage mmlu_pro 2,000 Uniform across categories Robust multi-task knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_benchmark.textquestion-answering10K<n<100K0 likes25 downloads4mo agoHugging Face26cosmosai471 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K5 likes24 downloads11mo agoHugging Face27PersianML /persian-general-knowledge Dataset Card for persian-gk (Persian General Knowledge) Dataset Summary persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training. Language: Persian (fa) Size: 5 897 conversations, 2–8 turns… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge.textquestion-answering1K<n<10K0 likes23 downloads2mo agoHugging Face28spacekat99 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K0 likes22 downloads4mo agoHugging Face29OpceanAI /sota-generaltexttext-generation100K<n<1M0 likes21 downloads4mo agoHugging Face30datasetter458 /bash-reference-manual-general-QAs Dataset generated from bash reference manual. book information like date and bash version are available within the very first rows of the dataset this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book columns : "Question", "Answer" texttext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.