datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TamilNadu-State-board-books-2025-Tamil-and-english-versions
Tamil Nadu School Textbooks — Tamil and English Structured Text
This dataset contains text extracted from 312 Tamil Nadu State Board school
textbooks for Standards 1–12. It covers Tamil- and English-medium books and
provides each retained book in two forms:
structured JSON with book metadata, ordered sections, typed content blocks,
source references, extraction statistics, and curation provenance;
Markdown for reading, inspection, and downstream text processing.
The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.olympiad-books-open-source
olympiad-books-open-source
Chunked content from 12 open-source mathematics textbooks, suitable for retrieval (RAG), embedding, and math reasoning research.
Source code: github.com/yoonholee/olympiad-books-open-source-pipeline
Books
Book
Author(s)
License
Source
An Infinitely Large Napkin
Evan Chen
CC BY-SA 4.0 / GPL v3
GitHub
Mathematical Reasoning: Writing and Proof
Ted Sundstrom
CC BY-NC-SA 3.0
GitHub
Exploring Combinatorial Mathematics
Richard Grassl… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/olympiad-books-open-source.aihub_mrc_books
Dataset Card for "mrc_aihub_books"
도서자료 기계독해
africans-history-books-qa-evalMedvik-Books
Medvik-Books - training dataset
Mappings of Authority main headings (the first author only) to the related book titles - based on Medvik system exports.
License
Medvik-Books - training dataset © 2025 by National Medical Library
is licensed under Creative Commons Attribution 4.0 International
Structure
"text1","text2","category"
"Author heading","Doc title","code1|code2"
category - multiple values-codes separated by a pipe… See the full description on the dataset page: https://huggingface.co/datasets/NLK-NML/Medvik-Books.Books-General-Linux
Linux Books Dataset
Dataset Description
The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants.
This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.booksort
Dataset Card for BookSORT
Dataset Description
Repository:
Paper: https://arxiv.org/abs/2410.08133
Point of Contact:
Dataset Summary
BookSORT is a dataset created from books for evaluation on the Sequence Order Recall Task (SORT), which assesses a model's ability to use temporal context in memory. SORT evaluation samples can be constructed from any sequential data. For BookSORT, the sequences are derived from text from 9 English language books that were… See the full description on the dataset page: https://huggingface.co/datasets/memari/booksort.books_question_answer
Dataset Card for books_question_answer
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Prarabdha/books_question_answer/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/books_question_answer.
