datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal_textbook
Multimodal-Textbook-6.5M
Overview
This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos.
It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts.
All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.TextbookReasoning
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Dataset Description
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.ultradata-math-textbook-exercise-ar
ultradata-math-textbook-exercise-ar
Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
cosmopedia-v2-textbook-and-howto-8.3m
Cosmopedia V2 Textbook and WikiHow Dataset 8.3M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.medical_textbooks_mcq
Medical Textbooks MCQs Dataset
This dataset is derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. It augments the original text snippets with synthetically generated Multiple Choice Questions (MCQs) in JSON format, suitable for fine-tuning or evaluating language models on medical MCQ generation tasks.
Dataset Details
Dataset Description
The source data consists of text snippets from the Textbooks corpus, a collection of 18 widely… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcq.cosmopedia-v2-textbook-and-howto-4.5m
Cosmopedia V2 Textbook and WikiHow Dataset 4.5M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-4.5m.sales-textbook-based
Dataset to train an online salesman model
This dataset was created for the purpose of training a sales agent chatbot that can convince people.
The initial idea came from: textbooks is all you need https://arxiv.org/abs/2306.11644
The original dataset is from https://huggingface.co/goendalf666
DeepSeek-V4-Flash (resoning: none) was used for the generation
Structure
textbook is just txt for pre-training
salesman_conversations is in sharedgpt format… See the full description on the dataset page: https://huggingface.co/datasets/loginik2/sales-textbook-based.Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tiny-strange-textbooks
Quirky Textbook Trove: Compact Excellence for Small Language Model
Strange dataset is 100% AI-generated, a compilation aligned with the vision of the Textbooks Are All You Need and Textbooks Are All You Need II: phi-1.5 technical report research. This dataset features 2,7M synthetic textbooks, encapsulating 16GB of raw text data. The unique name reflects its unconventional synthesis methodology, its compact size, deduped, and its emphasis on clear, focused content.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-strange-textbooks.Bangla-TextBook
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
---
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.sales-textbook_for_convincing_and_selling
Dataset Card for sales-textbook_for_convincing_and_selling
A textbook create for the purpose of training a sales chatbot.
Inspiration come from: Textbooks is all you need https://arxiv.org/abs/2306.11644
The data was generated by gpt-3.5-turbo
#Structure
A simpel textbook that has subheadlines and headlines.
Chapters and Subheadlines are mentioned in the dataset. Look at the first two examples.
Data Generation
The following code was used for the text generation:… See the full description on the dataset page: https://huggingface.co/datasets/goendalf666/sales-textbook_for_convincing_and_selling.tiny-code-textbooks
Code Explanation Textbooks
A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook.
openstax-pl-textbooks
OpenStax Poland academic textbooks
Text-only research contribution from eight Polish textbook volumes: physics
(three volumes), psychology, microeconomics, macroeconomics, marketing and nutrition.
Discovered pages: 1,612
Retained documents: 1,422
Tokens: 4,632,374 (cl100k_base proxy, measured on retained text)
Characters: 12,375,066
License: CC BY 4.0, documented separately in each preserved Polish foreword.
Snapshot payload commit: 5b31f74eb3733b330c6093dabc919fe304f6ef6f… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/openstax-pl-textbooks.nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.cosmopedia-v2-textbook-and-howto-2.3m
Cosmopedia V2 Textbook and WikiHow Dataset 2.3M
This dataset is derived from the HuggingFaceTB Smollm-Corpus with
a specific focus on the Cosmopedia V2 subset.
It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types.
Overview
The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for
researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-2.3m.korean-textbooks-edu
🇰🇷📚 korean-textbooks-edu
maywell/korean_textbooks의 모든 subset을 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋
불러오기
from datasets import load_dataset
ds = load_dataset("devngho/korean-textbooks-edu", name="scored_over_3", split="train")
성능
예정
컴퓨팅
Google Cloud TPU, transformers, JAX, tpuswarm
하드웨어
TPU v4-8 x 4 instances, 약 2시간 소요
이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡
라이선스
원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-textbooks-edu.codecontests-textbooks-dp-v1This dataset is a synthetic collection designed for algorithmic problem-solving, particularly in the dynamic programming domain. It is inspired by problems from the DeepMind/code_contests dataset, ensuring authenticity and relevance to competitive programming and algorithmic challenges.
The dataset includes detailed problem statements, input-output specifications, constraints, and illustrative test cases. Each example mirrors real-world scenarios, providing not only the problem but also… See the full description on the dataset page: https://huggingface.co/datasets/bblain/codecontests-textbooks-dp-v1.tiny-orca-textbooks
Textbook-like Dataset: A Comprehensive Resource for Text-Based Skills Development in Small Language Models
This dataset is a collection of 147k synthetic textbooks designed to enhance the text-based skills of small language models. The curriculum is meticulously structured to progress from simple to complex tasks, ensuring a gradual and effective learning experience during pretraining or finetuning SLMs.
The inspiration for this dataset comes from the technical report paper… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-orca-textbooks.textbook-qa-nepali-reasoning
Textbook Question-Answering Dataset (Nepali)
This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline.
Splits
train: validated conversations with non-empty question, answer, and rephrased_text.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/textbook-qa-nepali-reasoning")
train = ds["train"]
Schema
train: each row contains:
id: unique string
conversations: list of N messages (N ≥… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbook-qa-nepali-reasoning.textbooks-qa-nepali
Textbook Question-Answering Dataset (Nepali)
This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline.
Splits
train: validated conversations with non-empty question, answer, and rephrased_text.
Usage
from datasets import load_dataset
ds = load_dataset("dineshkarki/textbooks-qa-nepali")
train = ds["train"]
Schema
train: each row contains:
id: unique string
conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.socratic-cybernetics-textbook
De Balans-Code: Cybernetische Synthese van Hardware, Software en Kosmische Ankers
Het Geconsolideerde Systeem-Protocol — Candor-Modus Editie
Systeembasis: Decentrale Validatienetwerken ($TAO) vanaf 2026
Hoofdstuk 1: Systeemarchitectuur & Fysieke Infrastructuur
Socio-Economische Centralisatie, Cybernetische Balans en Empirische Analyse van Megalithische Logistiek
1.1 INLEIDING: SYSTEMISCHE ENTROPIE EN HET… See the full description on the dataset page: https://huggingface.co/datasets/Feedbackloop369/socratic-cybernetics-textbook.Malaysia-textbook-cleaned
Malaysia-textbook-cleaned
Cleaned by Kureiwa.
Cleaned version of Scicom-intl/Malaysia-Textbook, which gathers KSSR and KSSM textbooks in PDF format and converts them to text using Qwen/Qwen3-235B-A22B-Instruct-2507. Covers Bahasa Melayu, Chinese, English, Tamil, and Arabic/Jawi subjects.
Cleaning process
The cleaning pipeline (scripts/clean.py) is pure Python stdlib (csv, re), with DuckDB CLI used only for Parquet I/O. Each page's content is processed as follows:… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/Malaysia-textbook-cleaned.medical_textbooks_mcmq
Medical Textbooks French MCQ Fine-tuning Dataset
This dataset provides fine-tuning data derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. Using French text synthetically generated from the original English snippets, it aims to train models to answer medical Multiple Choice Questions (MCQs). Specifically, the model is presented with a JSON object containing the question and options, and it should generate a JSON object containing the correct options and… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcmq.swift-python-textbook-20260218textbooks-lite-700k-sharegpt-enPurified-openai-messages
📖 textbooks-lite-700k-enPurified-openai-messages
textbooks-lite-700k-enPurified is a highly curated, "prose-first" subset of the original jtatman/textbooks-lite-700k-sharegpt.
The enPurified collection is built on a specific philosophy: Specialization through Purity. While the ecosystem is rich with datasets for competitive programming and complex mathematics, high-quality, fluent English prose is often diluted by technical syntax or symbolic logic.
For this dataset, the enPurified… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/textbooks-lite-700k-sharegpt-enPurified-openai-messages.swift-python-textbook-20260219swift-python-textbook-20260302Spirit_Kings_Golden_Textbook
About
This is a dataset about the Spirit Kings clan from the Mineberry Minecraft server.
