CoolFace
24 results

ara

MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes36k downloads2y agoHugging FaceBangumiBase /arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu. This is the image base of bangumi A-Rank Party wo Ridatsu shita Ore wa, Moto Oshiego-tachi to Meikyuu Shinbu wo Mezasu., we detected 168 characters, 15579 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu.image10K<n<100K0 likes9.9k downloads1y agoHugging Facenabbahi /traditionnals_arabic_shoes_split3d1K<n<10K0 likes8.3k downloads4mo agoHugging Facearanemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.1k downloads3mo agoHugging FaceMBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.8k downloads2y agoHugging Faceoddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M28 likes3.7k downloads2mo agoHugging Face

People