datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.pdf-upload-caps-2026
PDF upload caps on public portals (2026)
How large can a PDF be before a government, university or job portal rejects it? This dataset records the published file size limit of 162 portals in France, the United States, the United Kingdom, Germany, Spain and India, each with the exact wording of the limit and a link to the official page where it was found.
It was collected in September 2026 for the EasyPDF study The 1 MB Problem: PDF File Size Statistics for 2026 (French version:… See the full description on the dataset page: https://huggingface.co/datasets/EasyPDF/pdf-upload-caps-2026.combined_leaderboard_with_pdf_scoresNNProject_embeddings_Example_pdf
