datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afrimmlu
Dataset Card for afrimmlu
Dataset Summary
AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages.
It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.MasrawiQA
MasrawiQA
MasrawiQA is a culturally rooted, manually curated evaluation benchmark containing 129 question-answer pairs designed exclusively for the Cairene and Lower Egyptian Arabic dialect (arz_arab). Introduced as part of the MRL 2026 Shared Task, this dataset evaluates language models on authentic street slang, everyday idioms, and dense regional knowledge spanning Egyptian sports history, local cuisine, and urban pop culture.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/JohnnyGeorge/MasrawiQA.Bible
Full Bible Chapter wise - Tamil
Web Scrapped from https://bible.catholicgallery.org/ecu-tamil/
