datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CryoLithe-training-datasetThe training Dataset for CryoLithe Models
The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058
Icecream reconstructions were obtained by splitting across angles.
Whenever available, we also provide dose-fractionated tilt series.
Dataset format:
Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles.
Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.SADA22
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.sada-train-preprocessedsada-train-wav2vec2-xls-r-300m-ar-preprocessedJitOPD-OpenR1-Math-220k-Teacher-Memory
JitOPD OpenR1-Math-220k Teacher Prefix-Logit Memory
This dataset contains sparse teacher next-token logits collected for JitOPD
retrieval-augmented decoding. The source prompts are the default configuration
of open-r1/OpenR1-Math-220k,
and the teacher is
Qwen/Qwen2.5-Math-7B-Instruct.
Only teacher trajectories whose final boxed answer passes both a numeric
signature prefilter and Math-Verify are retained. This release contains raw
teacher prefix/logit memory and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/sadadasdasdas/JitOPD-OpenR1-Math-220k-Teacher-Memory.sada_preprocessed_whisper_smallsada2022
Dataset Card for SADA - Saudi Audio Dataset for Arabic
Dataset Details
Dataset Description
The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.).
Curated by: The National… See the full description on the dataset page: https://huggingface.co/datasets/m6011/sada2022.CPMKsada22-najdisada2022
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.sada_clean_envSADA_khaledalganemsada2022_Rawdate
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.30k-SADA22_Saudisada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.wikipedia
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump, and you… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/wikipedia.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.sa-data
Storia dell'Arte Dataset (SA-Data)
📌 Descrizione del Dataset
Il dataset SA-Data è una raccolta strutturata di articoli della rivista Storia dell'Arte (https://www.storiadellarterivista.it/) digitalizzati e arricchiti con metadati dettagliati e rappresentazioni semantiche. È stato creato per supportare la ricerca accademica e le applicazioni di elaborazione del linguaggio naturale.
🔍 Contenuto
Il dataset include:
1050 articoli pubblicati tra il… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/sa-data.SADA_DATA
SADA - Saudi Audio Dataset for Arabic - Version 1.0
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), in collaboration with the Saudi Broadcasting Authority (SBA), published the “SADA” dataset, which stands for "Saudi Audio Dataset for Arabic”.
This dataset contains audio recordings sourced from more than 57 TV shows provided by the Saudi Broadcasting Authority. The total number of hours published for these recordings is… See the full description on the dataset page: https://huggingface.co/datasets/SassiiChaima/SADA_DATA.60H-SADA22-Saudisada_segmented_trainarabic-speech-SADA22-MSA
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.sadas
My Spraix Dataset
This dataset is a demo created for the Animator2D-v1.0.0 project, with the goal of generating pixel-art sprite animations from textual prompts. It is still in the preparation phase and contains sprite frames extracted and organized into subfolders, each with associated metadata.
Dataset Details
Format: PNG for frames, JSON for metadata.
Structure: Numbered subfolders (e.g., 152_frames) containing sequential frames (e.g., frame_000.png).
Usage: Designed… See the full description on the dataset page: https://huggingface.co/datasets/Loacky/sadas.sada-validation-wav2vec2-xls-r-300m-ar-preprocessedarabic-speech-SADA22-Khaliji
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.SADA-Najdi-from-kaggle
SADA 2022 — Najdi Dialect (Preprocessed for TTS)
Najdi dialect subset of SADA 2022,
preprocessed and ready for XTTS-v2 fine-tuning.
Preprocessing Pipeline
SADA full episodes → Filter Najdi → Slice by SegmentStart/End
→ Resample 22050Hz → Trim silence → Peak normalize
→ Duration filter (2.0–11.0s)
→ SNR filter (≥12.0dB)
→ Environment filter (Clean only)
→ Arabic text normalization
Step
Details
Dialect
SpeakerDialect == "Najdi"
Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.sadASDasdsadai-mrec-query-rewrite-13kbangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.arabic-eou-sada22
Arabic End-of-Utterance Detection Dataset
Dataset Description
This dataset is designed for training Arabic End-of-Utterance (EOU) detection models for real-time voice agents. It contains complete and incomplete Arabic utterances derived from the SADA22 dataset with emphasis on Saudi dialect.
Dataset Structure
Files
train.csv: 125,155 samples (62,578 complete, 62,577 incomplete)
test.csv: 31,289 samples (15,644 complete, 15,645 incomplete)… See the full description on the dataset page: https://huggingface.co/datasets/azeddinShr/arabic-eou-sada22.MSciNLI
MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference
This repository contains the dataset for the NAACL 2024 paper "MSCINLI: A Diverse Benchmark for Scientific Natural Language Inference."
If you face any difficulties while downloading the dataset, raise an issue in this repository or contact us at msadat3@uic.edu.
For more details about the dataset, please visit: https://github.com/msadat3/MSciNLI
Citation
If you use this dataset, please cite our… See the full description on the dataset page: https://huggingface.co/datasets/sadat2307/MSciNLI.
