datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.BLEnD-Vis
BLEnD-Vis
BLEnD-Vis is a benchmark for evaluating vision-language models (VLMs) on culturally grounded multiple-choice questions, including a text-only setting and a visual setting with generated images.
Paper: https://arxiv.org/abs/2510.11178
Dataset repo: https://huggingface.co/datasets/Incomple/BLEnD-Vis
Code: https://github.com/Social-AI-Studio/BLEnD-Vis
Source
BLEnD-Vis is derived from the BLEnD dataset on Hugging Face (nayeon212/BLEnD).
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Incomple/BLEnD-Vis.english_telugu_slangenglish_tamil_slang
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BLESSENA30/english_tamil_slang.Novachrono-Reasoning-Blend-v1
🧠 Novachrono-Reasoning-Blend-v1
Novachrono-Reasoning-Blend-v1 is a large-scale, multi-source instruction dataset designed for training and evaluating reasoning-capable language models. The dataset contains structured instructions, intermediate reasoning annotations, and high-quality final responses across a diverse range of tasks and domains.
Built with a strong emphasis on clarity, consistency, and practical usefulness, this dataset is intended for instruction tuning, alignment… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/Novachrono-Reasoning-Blend-v1.38_blessings_mingalar_tayartaw_qna
၃၈ ဖြာ မင်္ဂလာ တရားတော် [The 38 Blessings of Maha Mangala Sutta Q&A]
Created by: freococoLicense: CC0 1.0 Universal (Public Domain)
Summary
This dataset consists of 1,111 high-quality Questions and Answers centered on the 38 Blessings (၃၈ ဖြာ မင်္ဂလာ - Mangala). The answers are written in a Burmese Spoken Style to ensure natural AI conversational flow, while maintaining deep Philosophical, Psychological, Social, and Spiritual perspectives.
This repository is a specialized… See the full description on the dataset page: https://huggingface.co/datasets/freococo/38_blessings_mingalar_tayartaw_qna.bleta-sq-dataset-v1
Bleta SQ Instruct v1
Cleaned instruction-following dataset for Albanian language fine-tuning, used to train the Bleta AI assistant.
Dataset Details
Total rows: 39,873
Language: Albanian (sq)
Format: Alpaca (instruction / input / output)
Composition
Split
Rows
Description
Albanian Alpaca
38,480
Cleaned from saillab/alpaca-albanian-cleaned (removed ~12K Afrikaans rows)
Bleta Identity
1,393
Grammatically correct Albanian identity Q&A for the Bleta… See the full description on the dataset page: https://huggingface.co/datasets/klei1/bleta-sq-dataset-v1.
