CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DAMO-NLP-SG /multimodal_textbook Multimodal-Textbook-6.5M Overview This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos. It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts. All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.text-generation1M<n<10M164 likes5.3k downloads2y agoHugging Face02MegaScience /TextbookReasoning MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Dataset Description Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.texttext-generation100K<n<1M33 likes1.9k downloads1y agoHugging Face03nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes698 downloads2y agoHugging Face04SultanR /ultradata-math-textbook-exercise-ar ultradata-math-textbook-exercise-ar Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.texttext-generation10M<n<100M0 likes468 downloads1mo agoHugging Face05vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split UltraData-Math L3 Textbook Exercise Synthetic Split Source dataset: openbmb/UltraData-Math Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic Each row contains: uid question answer The original content field was split using the literal markers The exercise: and The solution:. texttext-generation10M<n<100M1 likes429 downloads6mo agoHugging Face06schuler /cosmopedia-v2-textbook-and-howto-8.3m Cosmopedia V2 Textbook and WikiHow Dataset 8.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.texttext-generation1M<n<10M5 likes300 downloads2y agoHugging Face07asanchez75 /medical_textbooks_mcq Medical Textbooks MCQs Dataset This dataset is derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. It augments the original text snippets with synthetically generated Multiple Choice Questions (MCQs) in JSON format, suitable for fine-tuning or evaluating language models on medical MCQ generation tasks. Dataset Details Dataset Description The source data consists of text snippets from the Textbooks corpus, a collection of 18 widely… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcq.textmultiple-choice1K<n<10K0 likes233 downloads1y agoHugging Face08schuler /cosmopedia-v2-textbook-and-howto-4.5m Cosmopedia V2 Textbook and WikiHow Dataset 4.5M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-4.5m.texttext-generation1M<n<10M0 likes213 downloads2y agoHugging Face09loginik2 /sales-textbook-based Dataset to train an online salesman model This dataset was created for the purpose of training a sales agent chatbot that can convince people. The initial idea came from: textbooks is all you need https://arxiv.org/abs/2306.11644 The original dataset is from https://huggingface.co/goendalf666 DeepSeek-V4-Flash (resoning: none) was used for the generation Structure textbook is just txt for pre-training salesman_conversations is in sharedgpt format… See the full description on the dataset page: https://huggingface.co/datasets/loginik2/sales-textbook-based.text-generation1K<n<10K0 likes175 downloads24d agoHugging Face10tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes173 downloads2mo agoHugging Face11nampdn-ai /tiny-strange-textbooksgated Quirky Textbook Trove: Compact Excellence for Small Language Model Strange dataset is 100% AI-generated, a compilation aligned with the vision of the Textbooks Are All You Need and Textbooks Are All You Need II: phi-1.5 technical report research. This dataset features 2,7M synthetic textbooks, encapsulating 16GB of raw text data. The unique name reflects its unconventional synthesis methodology, its compact size, deduped, and its emphasis on clear, focused content. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-strange-textbooks.texttext-generation1M<n<10M94 likes172 downloads3y agoHugging Face12md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes135 downloads1y agoHugging Face13goendalf666 /sales-textbook_for_convincing_and_selling Dataset Card for sales-textbook_for_convincing_and_selling A textbook create for the purpose of training a sales chatbot. Inspiration come from: Textbooks is all you need https://arxiv.org/abs/2306.11644 The data was generated by gpt-3.5-turbo #Structure A simpel textbook that has subheadlines and headlines. Chapters and Subheadlines are mentioned in the dataset. Look at the first two examples. Data Generation The following code was used for the text generation:… See the full description on the dataset page: https://huggingface.co/datasets/goendalf666/sales-textbook_for_convincing_and_selling.texttext-generation1K<n<10K24 likes120 downloads3y agoHugging Face14nampdn-ai /tiny-code-textbooksgated Code Explanation Textbooks A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook. tabulartext-generation100K<n<1M13 likes104 downloads3y agoHugging Face15PiotrSty /openstax-pl-textbooks OpenStax Poland academic textbooks Text-only research contribution from eight Polish textbook volumes: physics (three volumes), psychology, microeconomics, macroeconomics, marketing and nutrition. Discovered pages: 1,612 Retained documents: 1,422 Tokens: 4,632,374 (cl100k_base proxy, measured on retained text) Characters: 12,375,066 License: CC BY 4.0, documented separately in each preserved Polish foreword. Snapshot payload commit: 5b31f74eb3733b330c6093dabc919fe304f6ef6f… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/openstax-pl-textbooks.texttext-generation1K<n<10K0 likes72 downloads18d agoHugging Face16dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes63 downloads1y agoHugging Face17schuler /cosmopedia-v2-textbook-and-howto-2.3m Cosmopedia V2 Textbook and WikiHow Dataset 2.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-2.3m.texttext-generation1M<n<10M2 likes55 downloads2y agoHugging Face18devngho /korean-textbooks-edugated 🇰🇷📚 korean-textbooks-edu maywell/korean_textbooks의 모든 subset을 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋 불러오기 from datasets import load_dataset ds = load_dataset("devngho/korean-textbooks-edu", name="scored_over_3", split="train") 성능 예정 컴퓨팅 Google Cloud TPU, transformers, JAX, tpuswarm 하드웨어 TPU v4-8 x 4 instances, 약 2시간 소요 이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡ 라이선스 원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-textbooks-edu.texttext-generation10M<n<100M7 likes42 downloads2y agoHugging Face19bblain /codecontests-textbooks-dp-v1This dataset is a synthetic collection designed for algorithmic problem-solving, particularly in the dynamic programming domain. It is inspired by problems from the DeepMind/code_contests dataset, ensuring authenticity and relevance to competitive programming and algorithmic challenges. The dataset includes detailed problem statements, input-output specifications, constraints, and illustrative test cases. Each example mirrors real-world scenarios, providing not only the problem but also… See the full description on the dataset page: https://huggingface.co/datasets/bblain/codecontests-textbooks-dp-v1.texttext-generation1K<n<10K2 likes37 downloads2y agoHugging Face20nampdn-ai /tiny-orca-textbooksgated Textbook-like Dataset: A Comprehensive Resource for Text-Based Skills Development in Small Language Models This dataset is a collection of 147k synthetic textbooks designed to enhance the text-based skills of small language models. The curriculum is meticulously structured to progress from simple to complex tasks, ensuring a gradual and effective learning experience during pretraining or finetuning SLMs. The inspiration for this dataset comes from the technical report paper… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-orca-textbooks.texttext-generation100K<n<1M43 likes35 downloads3y agoHugging Face21dineshkarki /textbook-qa-nepali-reasoning Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbook-qa-nepali-reasoning") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of N messages (N ≥… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbook-qa-nepali-reasoning.textquestion-answering1K<n<10K1 likes32 downloads1y agoHugging Face22dineshkarki /textbooks-qa-nepali Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbooks-qa-nepali") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.tabularquestion-answering1K<n<10K1 likes30 downloads1y agoHugging Face23Feedbackloop369 /socratic-cybernetics-textbook De Balans-Code: Cybernetische Synthese van Hardware, Software en Kosmische Ankers Het Geconsolideerde Systeem-Protocol — Candor-Modus Editie Systeembasis: Decentrale Validatienetwerken ($TAO) vanaf 2026 Hoofdstuk 1: Systeemarchitectuur & Fysieke Infrastructuur Socio-Economische Centralisatie, Cybernetische Balans en Empirische Analyse van Megalithische Logistiek 1.1 INLEIDING: SYSTEMISCHE ENTROPIE EN HET… See the full description on the dataset page: https://huggingface.co/datasets/Feedbackloop369/socratic-cybernetics-textbook.text-generationn<1K0 likes29 downloads3d agoHugging Face24Kureiwa /Malaysia-textbook-cleaned Malaysia-textbook-cleaned Cleaned by Kureiwa. Cleaned version of Scicom-intl/Malaysia-Textbook, which gathers KSSR and KSSM textbooks in PDF format and converts them to text using Qwen/Qwen3-235B-A22B-Instruct-2507. Covers Bahasa Melayu, Chinese, English, Tamil, and Arabic/Jawi subjects. Cleaning process The cleaning pipeline (scripts/clean.py) is pure Python stdlib (csv, re), with DuckDB CLI used only for Parquet I/O. Each page's content is processed as follows:… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/Malaysia-textbook-cleaned.texttext-generation10K<n<100K0 likes25 downloads1mo agoHugging Face25asanchez75 /medical_textbooks_mcmq Medical Textbooks French MCQ Fine-tuning Dataset This dataset provides fine-tuning data derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. Using French text synthetically generated from the original English snippets, it aims to train models to answer medical Multiple Choice Questions (MCQs). Specifically, the model is presented with a JSON object containing the question and options, and it should generate a JSON object containing the correct options and… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcmq.textmultiple-choice1K<n<10K0 likes24 downloads1y agoHugging Face26igarin /swift-python-textbook-20260218texttext-generationn<1K0 likes24 downloads7mo agoHugging Face27enPurified /textbooks-lite-700k-sharegpt-enPurified-openai-messages 📖 textbooks-lite-700k-enPurified-openai-messages textbooks-lite-700k-enPurified is a highly curated, "prose-first" subset of the original jtatman/textbooks-lite-700k-sharegpt. The enPurified collection is built on a specific philosophy: Specialization through Purity. While the ecosystem is rich with datasets for competitive programming and complex mathematics, high-quality, fluent English prose is often diluted by technical syntax or symbolic logic. For this dataset, the enPurified… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/textbooks-lite-700k-sharegpt-enPurified-openai-messages.texttext-generation100K<n<1M1 likes22 downloads9mo agoHugging Face28igarin /swift-python-textbook-20260219texttext-generation1K<n<10K0 likes21 downloads7mo agoHugging Face29igarin /swift-python-textbook-20260302texttext-generation1K<n<10K0 likes20 downloads7mo agoHugging Face3016dvnk /Spirit_Kings_Golden_Textbook About This is a dataset about the Spirit Kings clan from the Mineberry Minecraft server. texttext-generation1K<n<10K1 likes16 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.