CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes612 downloads2mo agoHugging Face02tasal9 /ZamAi-Pashto-Datasets-V2 ZamAI Cleaned Pashto Dataset V2 Languages: psLicense: apache-2.0Task categories: summarization, text-generation, feature-extractionSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for summarization, text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAi-Pashto-Datasets-V2") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAi-Pashto-Datasets-V2.textsummarization10K<n<100K2 likes170 downloads2mo agoHugging Face03tasal9 /ZamAI-Pashto-Dataset-Cleaned ZamAI Pashto Dataset Cleaned Languages: psLicense: apache-2.0Task categories: text-classification, text-generation, question-answeringSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-classification, text-generation, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Dataset-Cleaned") print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Dataset-Cleaned.texttext-classification10K<n<100K0 likes115 downloads2mo agoHugging Face04tasal9 /ZamAI-Pashto-Mega-Dataset ZamAI Pashto Mega Dataset Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.texttext-generation1M<n<10M0 likes65 downloads2mo agoHugging Face05tasal9 /zamai-pashto-vision ZamAI Pashto Vision Languages: psLicense: cc-by-4.0Task categories: image-to-text, image-classificationSize categories: 1K<n<10K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for image-to-text, image-classification tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-vision") print(dataset) Configs pashto_captions: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-vision.imageimage-to-textn<1K0 likes41 downloads2mo agoHugging Face06tasal9 /ZamAI-Pashto-High-Qualituly-Dataset ZamAI Pashto High Quality Dataset Languages: psLicense: mitTask categories: text-generation, question-answering, translationSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, question-answering, translation tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-High-Qualituly-Dataset") print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-High-Qualituly-Dataset.texttext-generationn<1K0 likes40 downloads2mo agoHugging Face07ZamAI-Pashto /zamai-pashto-video ZamAI Pashto Video ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation. Dataset Summary The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows. Languages Pashto Dari English Modalities Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.textvideo-classificationn<1K0 likes39 downloads2mo agoHugging Face08ZamAI-Pashto /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest. Configs from datasets import load_dataset metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata") sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample") Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.tabularautomatic-speech-recognitionn<1K0 likes38 downloads2mo agoHugging Face09tasal9 /zamai-pashto-voice2voice ZamAI Pashto Voice2Voice Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-voice2voice") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.audioautomatic-speech-recognition1K<n<10K1 likes31 downloads2mo agoHugging Face10tasal9 /ZamAI-Pashto-MegaDataset-v1 ZamAI Pashto Mega Dataset v1 Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-MegaDataset-v1") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-MegaDataset-v1.texttext-generation10K<n<100K0 likes29 downloads2mo agoHugging Face11ZamAI-Pashto /zamai-pashto-vision ZamAI-Pashto Vision This repository contains the dataset and scripts for the ZamAI-Pashto vision project, focusing exclusively on Pashto-related visual data, including scene captions, signs, and cultural objects. Project Structure images/: Raw and cleaned images, including signs, boards, and documents with Pashto text. annotations/: Pashto captions, object bounding boxes, and scene labels. metadata/: Cultural and location-specific tags. scripts/: Image… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-vision.textn<1K0 likes28 downloads2mo agoHugging Face12ZamAI-Pashto /zamai-pashto-emotion ZamAI Pashto Emotion ZamAI Pashto Emotion is a dataset scaffold for emotional and cultural context modeling across Pashto-centered text, dialogue, and speech data. Dataset Summary The repository is organized for emotional language understanding, culturally grounded annotation, formality tracking, and speaker-aware metadata. Languages Pashto Dari English Modalities Text Dialogue Audio Project Structure… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-emotion.texttext-classificationn<1K0 likes26 downloads2mo agoHugging Face13tasal9 /zamai-pashto-documents Pashto Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-documents") print(dataset) Citation @misc{zamai_pashto_data, title = {{Pashto}}, author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes26 downloads2mo agoHugging Face14nassimjp /zamai-pashto-clean-cpt ZamAI Pashto Clean CPT This dataset is a hyper-cleaned, optimized, and fully deduplicated version of tasal9/ZamAI-Pashto-Mega-Dataset intended for Causal Language Modeling (CLM), pre-training, or fine-tuning text models in the Pashto language. 🛠️ Pipeline & Filtering Details Before processing the ~1.5 GB stream, an Internal Built-In Self-Test (BIST) was executed to verify environment I/O permissions, validate character ratio logic, and check connection… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/zamai-pashto-clean-cpt.text1M<n<10M0 likes26 downloads3mo agoHugging Face15tasal9 /ZamAI_Pashto_Datasetgated ZamAI Pashto Processed Dataset Languages: psLicense: cc-by-nd-4.0Task categories: text-generation, summarizationSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI_Pashto_Dataset") print(dataset) Configs instruction: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI_Pashto_Dataset.texttext-generation10K<n<100K0 likes9 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.