datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hani-ar-rifai-192kbps
Hani Ar-Rifai
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
hani-ar-rifai-192kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
192 kbps
Ayah files
6236 (2053 MiB)
Verified against the upstream MD5 list
6234
Ayahs absent upstream
0
Upstream folder
Hani_Rifai_192kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/hani-ar-rifai-192kbps.hani-ar-rifai-64kbps
Hani Ar-Rifai
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
hani-ar-rifai-64kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
64 kbps
Ayah files
6236 (702 MiB)
Verified against the upstream MD5 list
6234
Ayahs absent upstream
0
Upstream folder
Hani_Rifai_64kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is Al-Fatihah… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/hani-ar-rifai-64kbps.bengali-ocr-synthetic
Bengali OCR Synthetic Dataset
A high-quality synthetic Bengali OCR dataset for fine-tuning vision-language models like DeepSeek-OCR 2. Generated using 100+ professional Bengali Unicode fonts and 13K+ unique Bengali words with advanced text rendering via FreeType and HarfBuzz.
Dataset Overview
Language: Bengali (বাংলা)
Task: Optical Character Recognition (OCR)
Format: Conversation-based (vision-language)
Total Samples: 30,000
Train: 27,007 samples
Validation: 2,993… See the full description on the dataset page: https://huggingface.co/datasets/rifathridoy/bengali-ocr-synthetic.RIFA-Artbook-datasetRIFARSetupMOBILeditForensicULTRA9RIFAR3IDM_PCnabil-ar-rifai-48kbps
Nabil Ar-Rifai
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
nabil-ar-rifai-48kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
48 kbps
Ayah files
6345 (720 MiB)
Verified against the upstream MD5 list
0
Ayahs absent upstream
5
Upstream folder
Nabil_Rifa3i_48kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is Al-Fatihah… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/nabil-ar-rifai-48kbps.RIFAR2shikomori-stt-multi-v2Turing-Open-Reasoning
Computational STEM QA Dataset
Dataset Summary
This dataset contains computationally intensive, self-contained, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry.
Problems require multi-step reasoning, symbolic manipulation, numerical accuracy, or simulation-based verification. These tasks expose failure modes in state-of-the-art LLMs, making this dataset a strong benchmark for evaluating deep reasoning.
Each example… See the full description on the dataset page: https://huggingface.co/datasets/rifab988/Turing-Open-Reasoning.comorian_swadesh
Comorian Swadesh List
Dataset Summary
This dataset is a multilingual lexical resource linking words across four major languages and multiple Comorian regional varieties:
English (en)
French (fr)
Swahili (sw) and its normalized form (sw_clean)
Comorian language varieties:
swb — Shimaore (Mayotte)
wni — Shindzuani (Anjouan)
wlc — Shimwali (Moheli)
zdj — Shingazidja (Grande Comore)
The dataset supports linguistic comparison, cross-lingual NLP, and low-resource machine… See the full description on the dataset page: https://huggingface.co/datasets/rifailabs/comorian_swadesh.shimaore_dictionaryai-dubbing-videoshindzuani_lessonsshimwali_dictionaryquerydatasetshikomori_proverbsHeavy-Filesshikomori_phrasesshingazidja_dictionaryNoN_generic_248218_type_indian_drug_cleaned
Dataset Card for "NoN_generic_248218_type_indian_drug_cleaned"
More Information needed
shingazidja_sentences_jw
Dataset Card for "shingazidja-sentences-jw"
More Information needed
YGO-Turns
YGO-Turns: A Large-Scale Tournament Dataset of Expert Yu-Gi-Oh! Duel States
YGO-Turns is the first structured dataset of expert Yu-Gi-Oh! tournament play.
It contains 23,280 turn-level game state snapshots drawn from 1,792 duels
across 684 competitive matches played in online tournament settings.
This is the official dataset release accompanying the paper:
YGO-Turns: A Large-Scale Tournament Dataset of Expert Yu-Gi-Oh! Duel States
Rafael da Silva Santos, Leonardo Ferreira da… See the full description on the dataset page: https://huggingface.co/datasets/Rifa456/YGO-Turns.room_skechter_datasetsshimaore_lessonsroomskechter_floorplansshikomori-stt-multikmunlocker-ZKOS_POPSICLE_OS3
