datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zamai-pashto-documents
ZamAI-Pashto Documents
This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language.
Project Structure
data/: Contains scanned documents, extracted text, translations, and summaries.
annotations/: OCR bounding boxes, handwriting labels, and domain tags.
scripts/: OCR processing, text cleaning, and translation alignment scripts.
configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.ZamAi-Pashto-Datasets-V2
ZamAI Cleaned Pashto Dataset V2
Languages: psLicense: apache-2.0Task categories: summarization, text-generation, feature-extractionSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for summarization, text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAi-Pashto-Datasets-V2")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAi-Pashto-Datasets-V2.ZamAI-Pashto-Dataset-Cleaned
ZamAI Pashto Dataset Cleaned
Languages: psLicense: apache-2.0Task categories: text-classification, text-generation, question-answeringSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-classification, text-generation, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Dataset-Cleaned")
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Dataset-Cleaned.ZamAI-Pashto-Mega-Dataset
ZamAI Pashto Mega Dataset
Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.zamai-pashto-vision
ZamAI Pashto Vision
Languages: psLicense: cc-by-4.0Task categories: image-to-text, image-classificationSize categories: 1K<n<10K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for image-to-text, image-classification tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-vision")
print(dataset)
Configs
pashto_captions: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-vision.ZamAI-Pashto-High-Qualituly-Dataset
ZamAI Pashto High Quality Dataset
Languages: psLicense: mitTask categories: text-generation, question-answering, translationSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, question-answering, translation tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-High-Qualituly-Dataset")
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-High-Qualituly-Dataset.zamai-pashto-video
ZamAI Pashto Video
ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation.
Dataset Summary
The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows.
Languages
Pashto
Dari
English
Modalities
Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.ZamAI-Pashto-MegaDataset-v1
ZamAI Pashto Mega Dataset v1
Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-MegaDataset-v1")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-MegaDataset-v1.zamai-pashto-vision
ZamAI-Pashto Vision
This repository contains the dataset and scripts for the ZamAI-Pashto vision project, focusing exclusively on Pashto-related visual data, including scene captions, signs, and cultural objects.
Project Structure
images/: Raw and cleaned images, including signs, boards, and documents with Pashto text.
annotations/: Pashto captions, object bounding boxes, and scene labels.
metadata/: Cultural and location-specific tags.
scripts/: Image… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-vision.zamai-pashto-emotion
ZamAI Pashto Emotion
ZamAI Pashto Emotion is a dataset scaffold for emotional and cultural context modeling across Pashto-centered text, dialogue, and speech data.
Dataset Summary
The repository is organized for emotional language understanding, culturally grounded annotation, formality tracking, and speaker-aware metadata.
Languages
Pashto
Dari
English
Modalities
Text
Dialogue
Audio
Project Structure… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-emotion.zamai-pashto-documents
Pashto
Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-documents")
print(dataset)
Citation
@misc{zamai_pashto_data,
title = {{Pashto}},
author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.zamai-pashto-clean-cpt
ZamAI Pashto Clean CPT
This dataset is a hyper-cleaned, optimized, and fully deduplicated version of tasal9/ZamAI-Pashto-Mega-Dataset intended for Causal Language Modeling (CLM), pre-training, or fine-tuning text models in the Pashto language.
🛠️ Pipeline & Filtering Details
Before processing the ~1.5 GB stream, an Internal Built-In Self-Test (BIST) was executed to verify environment I/O permissions, validate character ratio logic, and check connection… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/zamai-pashto-clean-cpt.ZamAI_Pashto_Dataset
ZamAI Pashto Processed Dataset
Languages: psLicense: cc-by-nd-4.0Task categories: text-generation, summarizationSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI_Pashto_Dataset")
print(dataset)
Configs
instruction: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI_Pashto_Dataset.
