datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ot-lite
Open Telco Sample Data
1,850 telecom-specific evaluation samples across 8 benchmarks — designed for fast iteration during model development.
Use this dataset for fast iteration during model development. Evaluate against GSMA/ot-full for final results.
Eval Framework | Full Benchmarks
Benchmarks
| Config | Samples | Task | Paper |
|--------|--------:|------|-------|
| teleqna | 1,000 | Multiple-choice Q&A on telecom standards | arXiv |
| teletables | 100 | Table… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/ot-lite.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name: Huatuo26M-Lite
Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.XLRS-Bench-lite_VLM
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.XLRS-Bench-lite_VLM
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/zGinger/XLRS-Bench-lite_VLM.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.solveall-literature-priors
SolveAll Literature-Grounded Priors
Dataset summary
SolveAll Literature-Grounded Priors is an English-language dataset of
open-ended mathematical and scientific research problems paired with realistic
user priors whose epistemic relationship to the literature is explicitly
annotated. Each claim-bearing example is connected to one or more short
passages from identified literature sources. The passages are used to classify
the user's prior as contradicted, supported… See the full description on the dataset page: https://huggingface.co/datasets/PranathReddy/solveall-literature-priors.TIME-Lite
⌛️TIME-Lite: High-Quality Human-Annotated Subset for Temporal Reasoning Evaluation
🌐 Project Links
GitHub Repository: https://github.com/sylvain-wei/TIME
GitHub Project Page: https://omni-time.github.io
arXiv Paper: https://arxiv.org/pdf/2505.12891
TIME@HuggingFace: https://huggingface.co/datasets/SylvainWei/TIME
👋🏻 Introduction
⌛️TIME-Lite is a carefully curated human-annotated subset from the large-scale TIME benchmark dataset, containing 943… See the full description on the dataset page: https://huggingface.co/datasets/SylvainWei/TIME-Lite.mmmlu_lite
MMMLU-Lite
Introduction
A lite version of the MMMLU dataset, which is an community version of the MMMLU dataset by OpenCompass. Due to the large size of the original dataset (about 200k questions), we have created a lite version of the dataset to make it easier to use. We sample 25 examples from each language subject in the original dataset with fixed seed to ensure reproducibility, finally we have 19950 examples in the lite version of the dataset, which is about 10% of… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/mmmlu_lite.Global-MMLU-Lite
Global MMLU-Lite — Human Translated
Global MMLU-Lite is a multilingual
evaluation benchmark for LLMs covering 18 languages. This dataset extends it with professional
human translations for three additional low-resource languages that are not in the original:
Chichewa (nya), Māori (mri), and Inuktitut (iku).
Released as part of the BYOL: Bring Your Own Language Into LLMs
project (paper).
What's New
The original Global MMLU-Lite by Cohere
covers 18 languages: Arabic… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/Global-MMLU-Lite.global_mmlu_lite_pt
🌎 Global-MMLU Lite (Portuguese)
A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models
Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.literary-dataset-pack
Literary Dataset Pack
A rich and diverse multi-task instruction dataset generated from classic public domain literature.
📖 Overview
Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/dihdaidh/Huatuo26M-Lite.drosophila-literature-corpus
Drosophila Literature Corpus
A corpus of full-text scientific articles from PubMed Central related to Drosophila melanogaster gene research.
Dataset Description
This corpus contains ~17,000 full-text scientific articles downloaded from the PubMed Central Open Access Subset via the BioC-PMC API. The articles are associated with genes that have expert-curated summaries in FlyBase.
Usage
from datasets import load_dataset
corpus =… See the full description on the dataset page: https://huggingface.co/datasets/jimmyzxj/drosophila-literature-corpus.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/mingys/Huatuo26M-Lite.luganda-bilingual-literacy-exercises
Luganda-English Bilingual Literacy Exercises (P1–P3)
3,472 structured bilingual exercises for Ugandan primary school literacy instruction (Primary 1 through Primary 3). Each exercise contains parallel English and Luganda versions with questions, answers, and explanations.
Dataset Description
Grade
Exercises
File
P1
1,157
data/p1_exercises.json
P2
1,135
data/p2_exercises.json
P3
1,180
data/p3_exercises.json
Total
3,472
Exercise Types… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/luganda-bilingual-literacy-exercises.global_mmlu_lite_en
🌍 Global-MMLU Lite (English Only)
A Focused Benchmark for English-Language Reasoning in Large Language Models
Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_en.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name: Huatuo26M-Lite
Version:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuYingfeng/Huatuo26M-Lite.fiqh_doa_RAFT_ds_v01
Dataset Card — RAFT Islamic QA (Bilingual: Indonesia - Arab)
Dataset Retrieval-Augmented Fine-Tuning (RAFT) berbahasa Indonesia dan Arab (bersumber dari kitab Minhaj ath-Thalibin karya Imam An-Nawawi untuk Fiqh Syafii, serta himpunan Doa Harian dan Ibadah Praktis) yang dikembangkan oleh AI Literacy Innovation Institute (ALII) UIN SYARIF HIDAYATULLAH JAKARTA. Dataset ini dirancang khusus untuk melatih Large Language Models (LLM) agar dapat menjawab pertanyaan seputar hukum Islam… See the full description on the dataset page: https://huggingface.co/datasets/ai-literacy-innovation-institute/fiqh_doa_RAFT_ds_v01.global_mmlu_lite
🌍 Global-MMLU Lite Dataset
A Lightweight Benchmark for Multi-Domain Reasoning in Large Language Models
Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite.global_mmlu_lite_es
🌎 Global-MMLU Lite (Spanish Only)
A Focused Benchmark for Spanish-Language Reasoning in Large Language Models
Global-MMLU Lite (Spanish Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Spanish, providing a diverse and computationally efficient collection of fully translated and standardized QA samples across domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_es.RuFinQA-lite
RuFinQA — A Massive Multi-Task Reasoning Benchmark for Russian Financial Report Understanding
RuFinQA is a large-scale multi-task benchmark designed to evaluate the ability of language models to understand and reason over Russian statutory financial reports (Balance Sheet, Income Statement, Cash Flow Statement).
It contains 36,330 question–answer pairs across 5 task types, automatically derived from real-world corporate accounting statements obtained from open government data… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RuFinQA-lite.
