CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sada-group /CryoLithe-training-datasetThe training Dataset for CryoLithe Models The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058 Icecream reconstructions were obtained by splitting across angles. Whenever available, we also provide dose-fractionated tilt series. Dataset format: Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles. Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.fill-maskn<1K3 likes4.9k downloads3mo agoHugging Face02MohamedRashad /SADA22 Dataset Card for SADA (Saudi Audio Dataset for Arabic) Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.audioautomatic-speech-recognition100K<n<1M30 likes2.1k downloads1y agoHugging Face03mosama /sada-train-preprocessedaudio100K<n<1M0 likes605 downloads1y agoHugging Face04mosama /sada-train-wav2vec2-xls-r-300m-ar-preprocessedaudio100K<n<1M0 likes573 downloads1y agoHugging Face05sadadasdasdas /JitOPD-OpenR1-Math-220k-Teacher-Memory JitOPD OpenR1-Math-220k Teacher Prefix-Logit Memory This dataset contains sparse teacher next-token logits collected for JitOPD retrieval-augmented decoding. The source prompts are the default configuration of open-r1/OpenR1-Math-220k, and the teacher is Qwen/Qwen2.5-Math-7B-Instruct. Only teacher trajectories whose final boxed answer passes both a numeric signature prefilter and Math-Verify are retained. This release contains raw teacher prefix/logit memory and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/sadadasdasdas/JitOPD-OpenR1-Math-220k-Teacher-Memory.10K<n<100K0 likes300 downloads1mo agoHugging Face06mosama /sada_preprocessed_whisper_smallaudio100K<n<1M0 likes283 downloads1y agoHugging Face07m6011 /sada2022 Dataset Card for SADA - Saudi Audio Dataset for Arabic Dataset Details Dataset Description The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.). Curated by: The National… See the full description on the dataset page: https://huggingface.co/datasets/m6011/sada2022.audio1K<n<10K3 likes262 downloads2y agoHugging Face08sadadsss /CPMKimagen<1K2 likes255 downloads25d agoHugging Face09Abdalrahmankamel /sada22-najdiaudio10K<n<100K0 likes186 downloads1mo agoHugging Face10khaledalganem /sada2022 Dataset Card for SADA صدى Dataset Summary يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر. ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.audio100K<n<1M4 likes170 downloads2y agoHugging Face11mbuali /sada_clean_envtabular10K<n<100K0 likes170 downloads2y agoHugging Face12Sundus246 /SADA_khaledalganemsada2022_Rawdate Dataset Card for SADA صدى Dataset Summary يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر. ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.audio100K<n<1M0 likes129 downloads2mo agoHugging Face13shinnnsh23 /30k-SADA22_Saudiaudio10K<n<100K1 likes123 downloads2mo agoHugging Face14mosama /sada-validation-preprocessed Details This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo. In addtition, the following filters were applied to this data: All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.audioautomatic-speech-recognition1K<n<10K0 likes121 downloads1y agoHugging Face15Sadatsami /wikipedia Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump, and you… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/wikipedia.texttext-generation10M<n<100M0 likes117 downloads5mo agoHugging Face16aalshalfi /sada2022-arabic-tts SADA 2022 - Saudi Arabic Dataset for TTS مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech) المصدر الأصلي Kaggle: sdaiancai/sada2022 الاستخدام # طريقة 1: Git Clone !git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset # طريقة 2: مكتبة datasets from datasets import load_dataset dataset = load_dataset("aalshalfi/sada2022-arabic-tts") الملفات valid.csv - ملف البيانات الرئيسي wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.audio100K<n<1M0 likes114 downloads8mo agoHugging Face17paolodegasperis /sa-data Storia dell'Arte Dataset (SA-Data) 📌 Descrizione del Dataset Il dataset SA-Data è una raccolta strutturata di articoli della rivista Storia dell'Arte (https://www.storiadellarterivista.it/) digitalizzati e arricchiti con metadati dettagliati e rappresentazioni semantiche. È stato creato per supportare la ricerca accademica e le applicazioni di elaborazione del linguaggio naturale. 🔍 Contenuto Il dataset include: 1050 articoli pubblicati tra il… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/sa-data.tabulartoken-classification1K<n<10K1 likes113 downloads19d agoHugging Face18SassiiChaima /SADA_DATA SADA - Saudi Audio Dataset for Arabic - Version 1.0 The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), in collaboration with the Saudi Broadcasting Authority (SBA), published the “SADA” dataset, which stands for "Saudi Audio Dataset for Arabic”. This dataset contains audio recordings sourced from more than 57 TV shows provided by the Saudi Broadcasting Authority. The total number of hours published for these recordings is… See the full description on the dataset page: https://huggingface.co/datasets/SassiiChaima/SADA_DATA.audion<1K0 likes103 downloads1y agoHugging Face19MahmoudIbrahim /60H-SADA22-Saudiaudio10K<n<100K0 likes86 downloads8mo agoHugging Face20mosama /sada_segmented_train0 likes85 downloads1y agoHugging Face21badrex /arabic-speech-SADA22-MSA Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.audioautomatic-speech-recognition1K<n<10K2 likes83 downloads1y agoHugging Face22Loacky /sadas My Spraix Dataset This dataset is a demo created for the Animator2D-v1.0.0 project, with the goal of generating pixel-art sprite animations from textual prompts. It is still in the preparation phase and contains sprite frames extracted and organized into subfolders, each with associated metadata. Dataset Details Format: PNG for frames, JSON for metadata. Structure: Numbered subfolders (e.g., 152_frames) containing sequential frames (e.g., frame_000.png). Usage: Designed… See the full description on the dataset page: https://huggingface.co/datasets/Loacky/sadas.imagetext-classificationn<1K0 likes65 downloads1y agoHugging Face23mosama /sada-validation-wav2vec2-xls-r-300m-ar-preprocessedaudio1K<n<10K0 likes56 downloads1y agoHugging Face24badrex /arabic-speech-SADA22-Khaliji Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.audioautomatic-speech-recognition10K<n<100K4 likes52 downloads1y agoHugging Face25Ahmed007 /SADA-Najdi-from-kaggle SADA 2022 — Najdi Dialect (Preprocessed for TTS) Najdi dialect subset of SADA 2022, preprocessed and ready for XTTS-v2 fine-tuning. Preprocessing Pipeline SADA full episodes → Filter Najdi → Slice by SegmentStart/End → Resample 22050Hz → Trim silence → Peak normalize → Duration filter (2.0–11.0s) → SNR filter (≥12.0dB) → Environment filter (Clean only) → Arabic text normalization Step Details Dialect SpeakerDialect == "Najdi" Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.audiotext-to-speech10K<n<100K0 likes49 downloads5mo agoHugging Face26PERDYPTO /sadASDasd0 likes45 downloads2mo agoHugging Face27sadaisystems /sadai-mrec-query-rewrite-13ktabular10K<n<100K0 likes38 downloads1y agoHugging Face28Sadatsami /bangladesh-law-professional 🇧🇩 Bangladesh Law Professional Dataset A clean, instruction-tuned (Alpaca-style) question–answer dataset for fine-tuning language models on Bangladesh law, in Bangla and English. 👤 Author & Contribution Curated & built by Sadat Sami (@Sadatsami) Role Dataset architect — collected, cleaned, filtered, reformatted and published Motivation Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.textquestion-answering1K<n<10K0 likes37 downloads2mo agoHugging Face29azeddinShr /arabic-eou-sada22 Arabic End-of-Utterance Detection Dataset Dataset Description This dataset is designed for training Arabic End-of-Utterance (EOU) detection models for real-time voice agents. It contains complete and incomplete Arabic utterances derived from the SADA22 dataset with emphasis on Saudi dialect. Dataset Structure Files train.csv: 125,155 samples (62,578 complete, 62,577 incomplete) test.csv: 31,289 samples (15,644 complete, 15,645 incomplete)… See the full description on the dataset page: https://huggingface.co/datasets/azeddinShr/arabic-eou-sada22.text-classification100K<n<1M0 likes35 downloads9mo agoHugging Face30sadat2307 /MSciNLI MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference This repository contains the dataset for the NAACL 2024 paper "MSCINLI: A Diverse Benchmark for Scientific Natural Language Inference." If you face any difficulties while downloading the dataset, raise an issue in this repository or contact us at msadat3@uic.edu. For more details about the dataset, please visit: https://github.com/msadat3/MSciNLI Citation If you use this dataset, please cite our… See the full description on the dataset page: https://huggingface.co/datasets/sadat2307/MSciNLI.text100K<n<1M1 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.