datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_lmentry
Multi-LMentry
This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?".
Dataset Details
Dataset Description
Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集
Wizard-LM包含了很多难度超过Alpaca的指令。
中文的问题翻译会有少量指令注入导致翻译失败的情况
中文回答是根据中文问题再进行问询得到的。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.context
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.lm-similarity
Great Models Think Alike and this Undermines AI Oversight
This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight."
judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset.
judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.lme-gov
LME-Gov
LME-Gov is a project-maintained governance evaluation suite derived from
LongMemEval native histories. It tests state-transition behavior in writable
long-term memory systems: write admission, scoped retrieval, freshness
handling, contradiction/update handling, and leakage prevention.
This is the first public dataset release, LME-Gov v1.0
(dataset_version: 1.0.0). It is a reproducible benchmark artifact for
memory-governance research, not an independently maintained… See the full description on the dataset page: https://huggingface.co/datasets/siufgdaias/lme-gov.LMSYS-USP
LMSYS-USP Dataset
Overview
GitHub repository for exploring the source code and additional resources: https://github.com/wangkevin02/USP
The LMSYS-USP dataset contains high-quality dialogues with inferred user profiles(provide natural descriptions encompassing both objective facts and subjective characteristics), generated through a two-stage profiling pipeline (see our paper for details). The dataset includes a training set (87,882 examples), a validation set (4,626)… See the full description on the dataset page: https://huggingface.co/datasets/wangkevin02/LMSYS-USP.SlangBibTeX:
@article{liu2026lmlexiconimprovingdefinitionmodeling,
title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts},
author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan},
year={2026},
eprint={2602.14060},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.14060},
}
lmtq_999
🌏 中文多学科知识问答数据集 (Chinese Multi-disciplinary QA Dataset)
本数据集涵盖自然科学、人文社科、工程技术等多个维度的知识,旨在评估和提升模型在跨学科领域的推理与问答能力。数据集包含 多项选择题 (MCQ) 和 问答对 (QA) 两种形式,互为补充。
📊 数据集概览
数据集 ID: KaiLo2026/lmtq_999
总样本量: 999 条 (183 MCQ + 816 QA)
学科覆盖: 15+ 个主要领域 (天文、地学、生物、历史等)
语言: 简体中文 (zh-CN)
许可证: MIT License
适用任务: 知识问答、逻辑推理、学科能力评估、RAG 测试
📊 数据分布概览
1️⃣ MCQ 子集 (多项选择题)
总量: 183 条样本 | 特点: 适合评估模型的判别能力和知识广度。
学科领域
数量
占比
分布可视化 (比例缩放)
🌌 天文学
58
31.7%… See the full description on the dataset page: https://huggingface.co/datasets/KaiLo2026/lmtq_999.WordnetCitation:
@article{liu2026lmlexiconimprovingdefinitionmodeling,
title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts},
author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan},
year={2026},
eprint={2602.14060},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.14060},
}
Llama-3-70b-battlesChatbot Arena user conversations between Llama-3-70b VS GPT-4-1025 or Llama-3-70b VS Claude-3-Opus with user preference votes. Single turn. Excludes ties.
Used in Llama Data Analysis blog post and "VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models" (Paper, Code).
Citation
@article{dunlap_vibecheck,
title={VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models},
author={Lisa Dunlap and Krishna Mandal and Trevor… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Llama-3-70b-battles.multi_lmentry
Multi-LMentry
This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?".
Dataset Details
Dataset Description
Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/iperbole/multi_lmentry.VideoScienceBench
VideoScienceBench
A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Dataset Summary
Attribute
Value
Examples
160
Domains
Physics, Chemistry
Format
JSONL (prompt + expected phenomenon + vid)
Data Creation Pipeline
Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.WikiBibTeX:
@article{liu2026lmlexiconimprovingdefinitionmodeling,
title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts},
author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan},
year={2026},
eprint={2602.14060},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.14060},
}
