datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prophet-mosque-library
Prophet's Mosque Library
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.shamela-waqfeya-library
Shamela Waqfeya Library
📖 Overview
Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.prophet-mosque-library-compressed
Prophet's Mosque Library - Compressed
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.waqfeya-library-compressed
Waqfeya Library - Compressed
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.synonyms_dictionnaries
Description
Apache OpenOffice dictionnaries
shamela-waqfeya-library-compressed
Shamela Waqfeya Library - Compressed
📖 Overview
Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/shamela-waqfeya-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library-compressed.prompt-library
Qubax Prompt Library — 20 curated prompts
The public prompt library from the Qubax AI Studio — 20 curated, production-ready prompt templates across categories like writing, coding, analysis, and marketing.
Columns
Column
Description
name
Prompt title
category
Use-case category
prompt_template
Full prompt text
suggested_model
Model the prompt is tuned for
Notes
License: CC0 1.0 (public domain) — copy, remix, and redistribute… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/prompt-library.proxyquotes_library
The Proxy Quotes (pxyq) library
includes a fuction for calling the cell value with respect to a column and row of the csv dataset table. It calls for proxy stoploss distance, lotsize, and margins with leverages covering a betsize of 1 cash, commissions, swaps, spread, and more. It only have one simple function call, and that is pxyq.column('ASSET').
step 1: make sure you have pxyq.py in your directory. No need pip installations.
step 2: make an import pxyq is written on top of… See the full description on the dataset page: https://huggingface.co/datasets/algorembrant/proxyquotes_library.library-book-loans
Public Library Book Loans & Catalog Dataset (Free Sample)
This is a free sample with 3,503 rows. The full dataset has 36,653 rows across 4 tables.
Library circulation records for a simulated public library system with 3
branches, 12,000 catalog items, 5,000 patrons, and 20,000 loan transactions
over 2 years.
Features realistic patterns: summer reading program surge, academic year
peaks, genre popularity by age group, overdue rates, fine calculations,
and holds/reservations.… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/library-book-loans.exciting-library-21ed94
exciting-library-21ed94
Synthetic products test data: 40 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/OrbitRidge23/exciting-library-21ed94.library-dataStereotype-Elicitation-Prompt-LibraryLicense-Library
License Library
idk why i putted all this together in plaintext but here we are. literally spent way too much time hunting for clean plaintext official versions for projects so just dumped them here so u dont have to...
what is this?
Basically just folders with the LICENSE files inside. got MIT, Apache, and way too many GNU / CC variants. Tried to keep it clean...
also threw some metadata files (.csv, .json, .jsonl) in the metadata/ folder in case u wanna automate stuff… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/License-Library.wicopaco
Description
Voir https://wicopaco.limsi.fr/
Citation
@InProceedings{max10wicopaco,
author = {Aurélien Max and Guillaume Wisniewski},
title = {Mining Naturally-occurring Corrections and Paraphrases from
Wikipedia’s Revision History},
booktitle = {Proceedings of the Seventh conference on International
Language Resources and Evaluation (LREC'10)},
year = {2010},
month = {may},
date = {19-21},
address = {Valletta… See the full description on the dataset page: https://huggingface.co/datasets/fraug-library/wicopaco.LibraryFAQ1000national_library_of_korea_book_info
national_library_of_korea_book_info
국립중앙도서관에서 배포한, 국립중앙도서관에서 보관중인 도서 정보에 관한 데이터.
License
other (KOGL (Korea Open Government License) Type-1)
According to above KOGL, user can use public works freely and without fee regardless of its commercial use, and can change or modify to create secondary works when user complies with the terms provided as follows:
KOGL Type 1
Source Indication Liability
Users who use public works shall indicate source or copyright as follows:… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/national_library_of_korea_book_info.english_contractions_extensionswonef
Description
Voir https://wonef.pradet.me/documentation/
Citation
@inproceedings{pradet2014wonef,
title={WoNeF, an improved, expanded and evaluated automatic French translation of WordNet},
author={Pradet, Quentin and de Chalendar, Ga{\"e}l and Desormeaux Baguenier, Jeanne},
booktitle={Proceedings of the Seventh Global Wordnet Conference (GWC2014)},
pages={32--39},
year={2014},
month=jan
}
yale-library-entity-resolver-training-dataLibraryFAQDrug-like-Compound-Library
Molport Drug-Like Compound Library
The Molport Drug-Like Compound Library contains 2,140,604 pre-filtered drug-like compounds selected from Molport’s 5.8M purchasable compound database. This curated library is designed to support computational chemistry, medicinal chemistry, and drug discovery workflows, saving researchers time by focusing only on compounds that comply with the Lipinski Rule of 5 (Ro5) and additional filters.
Problem
Public chemical databases often… See the full description on the dataset page: https://huggingface.co/datasets/molport/Drug-like-Compound-Library.Library_FAQ_with_Context
