CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MegaScience /TextbookReasoning MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Dataset Description Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.texttext-generation100K<n<1M33 likes1.9k downloads1y agoHugging Face02nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes650 downloads2y agoHugging Face03SultanR /ultradata-math-textbook-exercise-ar ultradata-math-textbook-exercise-ar Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.texttext-generation10M<n<100M0 likes513 downloads1mo agoHugging Face04vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split UltraData-Math L3 Textbook Exercise Synthetic Split Source dataset: openbmb/UltraData-Math Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic Each row contains: uid question answer The original content field was split using the literal markers The exercise: and The solution:. texttext-generation10M<n<100M1 likes486 downloads6mo agoHugging Face05schuler /cosmopedia-v2-textbook-and-howto-8.3m Cosmopedia V2 Textbook and WikiHow Dataset 8.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-8.3m.texttext-generation1M<n<10M5 likes316 downloads2y agoHugging Face06asanchez75 /medical_textbooks_mcq Medical Textbooks MCQs Dataset This dataset is derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. It augments the original text snippets with synthetically generated Multiple Choice Questions (MCQs) in JSON format, suitable for fine-tuning or evaluating language models on medical MCQ generation tasks. Dataset Details Dataset Description The source data consists of text snippets from the Textbooks corpus, a collection of 18 widely… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcq.textmultiple-choice1K<n<10K0 likes233 downloads1y agoHugging Face07schuler /cosmopedia-v2-textbook-and-howto-4.5m Cosmopedia V2 Textbook and WikiHow Dataset 4.5M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-4.5m.texttext-generation1M<n<10M0 likes221 downloads2y agoHugging Face08nampdn-ai /tiny-strange-textbooksgated Quirky Textbook Trove: Compact Excellence for Small Language Model Strange dataset is 100% AI-generated, a compilation aligned with the vision of the Textbooks Are All You Need and Textbooks Are All You Need II: phi-1.5 technical report research. This dataset features 2,7M synthetic textbooks, encapsulating 16GB of raw text data. The unique name reflects its unconventional synthesis methodology, its compact size, deduped, and its emphasis on clear, focused content. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-strange-textbooks.texttext-generation1M<n<10M94 likes174 downloads3y agoHugging Face09tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes159 downloads2mo agoHugging Face10md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes144 downloads1y agoHugging Face11goendalf666 /sales-textbook_for_convincing_and_selling Dataset Card for sales-textbook_for_convincing_and_selling A textbook create for the purpose of training a sales chatbot. Inspiration come from: Textbooks is all you need https://arxiv.org/abs/2306.11644 The data was generated by gpt-3.5-turbo #Structure A simpel textbook that has subheadlines and headlines. Chapters and Subheadlines are mentioned in the dataset. Look at the first two examples. Data Generation The following code was used for the text generation:… See the full description on the dataset page: https://huggingface.co/datasets/goendalf666/sales-textbook_for_convincing_and_selling.texttext-generation1K<n<10K24 likes120 downloads3y agoHugging Face12nampdn-ai /tiny-code-textbooksgated Code Explanation Textbooks A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook. tabulartext-generation100K<n<1M13 likes104 downloads3y agoHugging Face13PiotrSty /openstax-pl-textbooks OpenStax Poland academic textbooks Text-only research contribution from eight Polish textbook volumes: physics (three volumes), psychology, microeconomics, macroeconomics, marketing and nutrition. Discovered pages: 1,612 Retained documents: 1,422 Tokens: 4,632,374 (cl100k_base proxy, measured on retained text) Characters: 12,375,066 License: CC BY 4.0, documented separately in each preserved Polish foreword. Snapshot payload commit: 5b31f74eb3733b330c6093dabc919fe304f6ef6f… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/openstax-pl-textbooks.texttext-generation1K<n<10K0 likes73 downloads20d agoHugging Face14schuler /cosmopedia-v2-textbook-and-howto-2.3m Cosmopedia V2 Textbook and WikiHow Dataset 2.3M This dataset is derived from the HuggingFaceTB Smollm-Corpus with a specific focus on the Cosmopedia V2 subset. It contains only entries that are categorized as either textbook, textbook_unconditionned_topic or WikiHow types. Overview The Cosmopedia Textbook and WikiHow Dataset is a collection of rows filtered from the original Smollm-Corpus dataset. This dataset is tailored for researchers and developers who require… See the full description on the dataset page: https://huggingface.co/datasets/schuler/cosmopedia-v2-textbook-and-howto-2.3m.texttext-generation1M<n<10M2 likes55 downloads2y agoHugging Face15dineshkarki /nepali-textbooks-corpus Nepali Textbooks Corpus for Grades 1-12 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 5634 Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12] Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.tabulartext-generation1K<n<10K2 likes43 downloads1y agoHugging Face16devngho /korean-textbooks-edugated 🇰🇷📚 korean-textbooks-edu maywell/korean_textbooks의 모든 subset을 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋 불러오기 from datasets import load_dataset ds = load_dataset("devngho/korean-textbooks-edu", name="scored_over_3", split="train") 성능 예정 컴퓨팅 Google Cloud TPU, transformers, JAX, tpuswarm 하드웨어 TPU v4-8 x 4 instances, 약 2시간 소요 이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡ 라이선스 원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-textbooks-edu.texttext-generation10M<n<100M7 likes40 downloads2y agoHugging Face17bblain /codecontests-textbooks-dp-v1This dataset is a synthetic collection designed for algorithmic problem-solving, particularly in the dynamic programming domain. It is inspired by problems from the DeepMind/code_contests dataset, ensuring authenticity and relevance to competitive programming and algorithmic challenges. The dataset includes detailed problem statements, input-output specifications, constraints, and illustrative test cases. Each example mirrors real-world scenarios, providing not only the problem but also… See the full description on the dataset page: https://huggingface.co/datasets/bblain/codecontests-textbooks-dp-v1.texttext-generation1K<n<10K2 likes39 downloads2y agoHugging Face18nampdn-ai /tiny-orca-textbooksgated Textbook-like Dataset: A Comprehensive Resource for Text-Based Skills Development in Small Language Models This dataset is a collection of 147k synthetic textbooks designed to enhance the text-based skills of small language models. The curriculum is meticulously structured to progress from simple to complex tasks, ensuring a gradual and effective learning experience during pretraining or finetuning SLMs. The inspiration for this dataset comes from the technical report paper… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-orca-textbooks.texttext-generation100K<n<1M43 likes36 downloads3y agoHugging Face19Kureiwa /Malaysia-textbook-cleaned Malaysia-textbook-cleaned Cleaned by Kureiwa. Cleaned version of Scicom-intl/Malaysia-Textbook, which gathers KSSR and KSSM textbooks in PDF format and converts them to text using Qwen/Qwen3-235B-A22B-Instruct-2507. Covers Bahasa Melayu, Chinese, English, Tamil, and Arabic/Jawi subjects. Cleaning process The cleaning pipeline (scripts/clean.py) is pure Python stdlib (csv, re), with DuckDB CLI used only for Parquet I/O. Each page's content is processed as follows:… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/Malaysia-textbook-cleaned.texttext-generation10K<n<100K0 likes26 downloads1mo agoHugging Face20asanchez75 /medical_textbooks_mcmq Medical Textbooks French MCQ Fine-tuning Dataset This dataset provides fine-tuning data derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. Using French text synthetically generated from the original English snippets, it aims to train models to answer medical Multiple Choice Questions (MCQs). Specifically, the model is presented with a JSON object containing the question and options, and it should generate a JSON object containing the correct options and… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcmq.textmultiple-choice1K<n<10K0 likes25 downloads1y agoHugging Face21enPurified /textbooks-lite-700k-sharegpt-enPurified-openai-messages 📖 textbooks-lite-700k-enPurified-openai-messages textbooks-lite-700k-enPurified is a highly curated, "prose-first" subset of the original jtatman/textbooks-lite-700k-sharegpt. The enPurified collection is built on a specific philosophy: Specialization through Purity. While the ecosystem is rich with datasets for competitive programming and complex mathematics, high-quality, fluent English prose is often diluted by technical syntax or symbolic logic. For this dataset, the enPurified… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/textbooks-lite-700k-sharegpt-enPurified-openai-messages.texttext-generation100K<n<1M1 likes24 downloads9mo agoHugging Face22dineshkarki /textbooks-qa-nepali Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbooks-qa-nepali") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of 2 messages: human and gpt… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbooks-qa-nepali.tabularquestion-answering1K<n<10K1 likes23 downloads1y agoHugging Face23igarin /swift-python-textbook-20260219texttext-generation1K<n<10K0 likes21 downloads7mo agoHugging Face24dineshkarki /textbook-qa-nepali-reasoning Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbook-qa-nepali-reasoning") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of N messages (N ≥… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbook-qa-nepali-reasoning.textquestion-answering1K<n<10K1 likes20 downloads1y agoHugging Face25igarin /swift-python-textbook-20260302texttext-generation1K<n<10K0 likes20 downloads7mo agoHugging Face26igarin /swift-python-textbook-20260218texttext-generationn<1K0 likes18 downloads7mo agoHugging Face2716dvnk /Spirit_Kings_Golden_Textbook About This is a dataset about the Spirit Kings clan from the Mineberry Minecraft server. texttext-generation1K<n<10K1 likes16 downloads11mo agoHugging Face28BadarHossain /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/BadarHossain/Bangla-TextBook.texttext-generation10K<n<100K0 likes15 downloads10mo agoHugging Face29InfoBayAI /Kannada-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Kannada STEM textbook data, containing 127 books and 6.64 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Kannada. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-STEM-Textbook-Dataset.tabulartext-generation10K<n<100K0 likes14 downloads9d agoHugging Face30dineshkarki /textbook-qa-nepali-multiturn Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbook-qa-nepali-multiturn") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of N messages (N ≥… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbook-qa-nepali-multiturn.textquestion-answering1K<n<10K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.